Back to feed
arXiv cs.CL
arXiv cs.CL
7/2/2026
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

Short summary

Knowledge-Based VQA benchmarks have systematic flaws—missing or contradicted answers, underspecified questions, and visually trivial scenes—that distort model rankings and overestimate VLM reasoning. Researchers propose audit-repair and augmentation protocols to fix benchmarks. Re-evaluation under corrected settings yields markedly different performance trends.

  • Existing KB-VQA benchmarks have systematic flaws: missing answers, unclear questions, oversimplified visual scenes
  • These flaws cause model ranking distortions and overestimate reasoning capabilities
  • Paper proposes audit-repair and multi-entity augmentation protocols to restore benchmark reliability

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more