Back to feed
AR
arXiv CS.AI
7/20/2026
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

Short summary

DrawingVQA is the first benchmark evaluating multimodal LLMs on real-world construction drawings, combining abstract geometry, symbolic notation, and domain-specific text. It includes 33 Issued-for-Construction drawings with 92 expert-curated QA pairs across three reasoning depths. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, especially at higher reasoning depths, highlighting the need for domain-specialized multimodal reasoning.

  • First benchmark for MLLM evaluation on real construction drawings with 33 drawings and 92 QA pairs
  • Three reasoning depths tested: perceptual, contextual, and domain-expert
  • SOTA MLLMs show significant performance gap vs human experts at higher reasoning depths

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more