Back to feed
arXiv cs.CL
arXiv cs.CL
7/14/2026
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

Short summary

CLIR-Bench is a new benchmark for question answering over irregular clinical time series, built from de-identified ICU records. It contains 6,600 QA instances across 11 clinical variables and 11 tasks, with each question linked to explicit temporal evidence. Experiments show existing generalist models struggle to retrieve and reason over sparse clinical observations, underscoring the need for better irregular time-series reasoning methods.

  • CLIR-Bench provides 6,600 QA instances over irregular clinical time series from ICU records
  • Each question links to explicit temporal evidence enabling evaluation of both accuracy and evidence use
  • Generalist models struggle with sparse clinical evidence, highlighting gaps in temporal reasoning

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more