arXiv cs.LG
6/30/2026

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models
Short summary
SciDraw-Bench is a new research benchmark designed to evaluate how well text-to-image and multimodal AI models can generate scientific figures like mechanism diagrams and experimental schematics. The benchmark introduces a four-dimensional evaluation protocol assessing text fidelity, semantic correctness, structural quality, and convention adherence across 32 tasks spanning 8 figure types and 10 scientific disciplines. Early results show domain-specific systems substantially outperform general-purpose models, with text fidelity remaining the most challenging dimension.
- •Introduces SciDraw-Bench: first benchmark for evaluating scientific figure generation
- •Four-dimensional evaluation protocol: text fidelity, semantic correctness, structure, convention adherence
- •Domain-specific models substantially outperform general text-to-image systems
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


