Back to feed
AR
arXiv CS.AI
6/30/2026
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

Short summary

IMCBench introduces a benchmark for evaluating eight multimodal LLMs in medical conversations, assessing safety, accuracy, and uncertainty quantification. Claude Opus 4.6 achieved the highest score (3.61/5), followed by Claude Sonnet 4.6 and GPT-5.2, though no model excelled uniformly. The research reveals visual inputs and EHR context are essential for safe clinical guidance, contradicting the assumption that diagnostic accuracy alone ensures patient safety.

  • Claude Opus 4.6 leads with 3.61/5; Claude Sonnet 4.6 and GPT-5.2 follow closely in multi-modal medical evaluation
  • No model dominates all dimensions; safety significantly degrades for malignant and rare medical conditions
  • Visual inputs and EHR context reduce safety drops by 0.18-0.23 on average when removed, proving both are critical

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more