Back to feed
arXiv cs.CL
arXiv cs.CL
7/17/2026
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Short summary

Researchers introduce Just Keep Prompting (JKP), a multi-turn evaluation framework testing VLM stability when users repeatedly challenge answers across up to 10 turns. GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B were evaluated on the STAR benchmark across 720 runs. While aggregate accuracy changes modestly, trajectory analysis reveals significant instability: correct answers regress, wrong answers recover, and repeated prompting often destabilizes rather than aids reasoning. Qwen3-VL-30B achieves highest final accuracy but becomes confidently wrong under contradiction; Gemini is stable but token-expensive; GPT-4o is most brittle and oscillatory.

  • JKP framework tests VLM stability across up to 10 turns of Socratic challenge
  • GPT-4o is most brittle; Gemini 2.5 Pro most stable but costly; Qwen3-VL-30B highest accuracy but confidently wrong under contradiction
  • Repeated prompting acts as destabilizer rather than reasoning aid, with strong model-dependent effects

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more