arXiv cs.CL
7/2/2026

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Short summary
Knowledge-Based VQA benchmarks have systematic flaws—missing or contradicted answers, underspecified questions, and visually trivial scenes—that distort model rankings and overestimate VLM reasoning. Researchers propose audit-repair and augmentation protocols to fix benchmarks. Re-evaluation under corrected settings yields markedly different performance trends.
- •Existing KB-VQA benchmarks have systematic flaws: missing answers, unclear questions, oversimplified visual scenes
- •These flaws cause model ranking distortions and overestimate reasoning capabilities
- •Paper proposes audit-repair and multi-entity augmentation protocols to restore benchmark reliability
Generated with AI, which can make mistakes.
Is this a good recommendation for you?