Back to feed
Dev.to
Dev.to
7/13/2026
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Short summary

A study evaluating nine vision-language models from 2017–2025 on a new Complex Social Behavior dataset of movie frames requiring multi-person interaction reasoning. Pre-MLLM captioners that looked strong on MS-COCO collapse on complex scenes, while modern MLLMs close most of the gap and eliminate four of five error types. The stubborn residual failure is spatial dependence reasoning, which persists even in the latest models.

  • New CSB benchmark of 100 movie frames exposes gaps hidden by MS-COCO's easy images
  • Modern MLLMs eliminate four of five visual-cognitive error types vs pre-MLLM captioners
  • Spatial dependence reasoning remains the most stubborn residual failure mode

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more