arXiv cs.CL
7/14/2026

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
Short summary
CLIR-Bench is a new benchmark for question answering over irregular clinical time series, built from de-identified ICU records. It contains 6,600 QA instances across 11 clinical variables and 11 tasks, with each question linked to explicit temporal evidence. Experiments show existing generalist models struggle to retrieve and reason over sparse clinical observations, underscoring the need for better irregular time-series reasoning methods.
- •CLIR-Bench provides 6,600 QA instances over irregular clinical time series from ICU records
- •Each question links to explicit temporal evidence enabling evaluation of both accuracy and evidence use
- •Generalist models struggle with sparse clinical evidence, highlighting gaps in temporal reasoning
Generated with AI, which can make mistakes.
Is this a good recommendation for you?