
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
Short summary
Researchers introduce Just Keep Prompting (JKP), a multi-turn evaluation framework testing VLM stability when users repeatedly challenge answers across up to 10 turns. GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B were evaluated on the STAR benchmark across 720 runs. While aggregate accuracy changes modestly, trajectory analysis reveals significant instability: correct answers regress, wrong answers recover, and repeated prompting often destabilizes rather than aids reasoning. Qwen3-VL-30B achieves highest final accuracy but becomes confidently wrong under contradiction; Gemini is stable but token-expensive; GPT-4o is most brittle and oscillatory.
- •JKP framework tests VLM stability across up to 10 turns of Socratic challenge
- •GPT-4o is most brittle; Gemini 2.5 Pro most stable but costly; Qwen3-VL-30B highest accuracy but confidently wrong under contradiction
- •Repeated prompting acts as destabilizer rather than reasoning aid, with strong model-dependent effects
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
