AR
arXiv CS.AI
7/20/2026

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Short summary
DrawingVQA is the first benchmark evaluating multimodal LLMs on real-world construction drawings, combining abstract geometry, symbolic notation, and domain-specific text. It includes 33 Issued-for-Construction drawings with 92 expert-curated QA pairs across three reasoning depths. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, especially at higher reasoning depths, highlighting the need for domain-specialized multimodal reasoning.
- •First benchmark for MLLM evaluation on real construction drawings with 33 drawings and 92 QA pairs
- •Three reasoning depths tested: perceptual, contextual, and domain-expert
- •SOTA MLLMs show significant performance gap vs human experts at higher reasoning depths
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

