AR
arXiv CS.AI
7/31/2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science
Short summary
ClinLens is a benchmark of 200 executable tasks across five linked MIMIC resources spanning EHRs, notes, ECGs, chest radiographs, and echocardiograms for evaluating clinical data-science agents. The best of 24 model-scaffold configurations achieves only 56.3% strict pass rate despite 100% execution success, while biomedical systems adapted to GPT-4o-mini reach at most 2.9%. Results expose a substantial gap between runnable code submissions and correct clinical analyses.
- •ClinLens benchmark covers 200 tasks across 5 MIMIC resources with structured EHR, notes, ECGs, radiographs, and echocardiograms
- •Best agent achieves 56.3% strict pass despite 100% execution success, revealing gap between runnable and correct clinical analyses
- •Biomedical systems adapted to GPT-4o-mini reach only 2.9% strict pass rate
Generated with AI, which can make mistakes.
Is this a good recommendation for you?