Back to feed
AR
arXiv CS.AI
6/25/2026
Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

Short summary

Researchers adapted classic human visual-search experiments to test whether vision-language models exhibit the same behavioral patterns, using reasoning token counts as a proxy for search effort. VLMs reproduce several human signatures (flat effort for feature search, climbing effort for conjunctions) but diverge: they expend more effort on target-absent trials and maintain perfect enumeration accuracy. The study demonstrates that psychophysical paradigms can probe machine visual cognition.

  • VLMs largely mirror human visual-search behavior, with reasoning tokens correlating to search effort
  • Key divergence: models expend more effort on target-absent trials, opposite to human patterns
  • Psychophysical paradigms offer an inexpensive method to probe and compare machine vs. human visual cognition

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more