arXiv cs.CL
7/20/2026

Large Language Models as Unified Multimodal Learners for Clinical Prediction
Short summary
This arXiv paper proposes converting all patient data—regardless of modality—into natural language sequences and fine-tuning a pretrained LLM end-to-end, eliminating task-specific fusion architectures. Evaluated across in-hospital mortality (MIMIC-III), graft failure prediction, and emergency triage, the unified textual serialization approach matches or exceeds specialized multimodal baselines and outperforms a clinically deployed gradient boosting model on graft failure. Results suggest bespoke fusion architectures are unnecessary for multimodal clinical prediction.
- •Converting all EHR data to text sequences and fine-tuning LLMs matches or beats task-specific multimodal fusion architectures
- •Tested on three clinical tasks: in-hospital mortality, graft failure, and emergency triage
- •Outperformed a clinically deployed gradient boosting system on graft failure prediction
Generated with AI, which can make mistakes.
Is this a good recommendation for you?