AR
arXiv CS.AI
6/30/2026

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
Short summary
IMCBench introduces a benchmark for evaluating eight multimodal LLMs in medical conversations, assessing safety, accuracy, and uncertainty quantification. Claude Opus 4.6 achieved the highest score (3.61/5), followed by Claude Sonnet 4.6 and GPT-5.2, though no model excelled uniformly. The research reveals visual inputs and EHR context are essential for safe clinical guidance, contradicting the assumption that diagnostic accuracy alone ensures patient safety.
- •Claude Opus 4.6 leads with 3.61/5; Claude Sonnet 4.6 and GPT-5.2 follow closely in multi-modal medical evaluation
- •No model dominates all dimensions; safety significantly degrades for malignant and rare medical conditions
- •Visual inputs and EHR context reduce safety drops by 0.18-0.23 on average when removed, proving both are critical
Generated with AI, which can make mistakes.
Is this a good recommendation for you?