AR
arXiv CS.AI
7/20/2026

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Short summary
This study decomposes which components of a coding agent—executable world models, simplification, and verification—drive performance on ARC-AGI-3. Testing four nested Codex-based agents across multiple GPT-5 variants, it finds that stronger models and greater reasoning effort consistently help, while individual component effects vary by setting. The full verification treatment ranks first everywhere and with gpt-5.6-sol solves every public game, though this likely reflects saturation of the public set only.
- •Stronger models and higher reasoning effort improve all agent variants on ARC-AGI-3
- •Verification treatment ranks first across all settings but uses substantially more resources
- •gpt-5.6-sol verification variant solves all public games, suggesting public-set saturation
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
