Back to feed
AR
arXiv CS.AI
7/13/2026
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Short summary

MedRealMM is a large-scale multimodal benchmark built from real de-identified patient-doctor interactions across a Chinese internet hospital, spanning 5,620 cases and 64 clinical departments. Results across 19 LLMs show that image information is critical for clinical performance and current frontier models still fall below physician quality, particularly in avoiding safety-sensitive errors. The dataset is publicly available on Hugging Face.

  • MedRealMM benchmarks 19 LLMs on 5,620 real-world multimodal medical consultation cases from a Chinese internet hospital
  • Image context is critical; frontier models still underperform physicians, especially in safety-sensitive error avoidance
  • Dataset is publicly available on Hugging Face for reproducible evaluation

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more