Back to feed
AR
arXiv CS.AI
7/20/2026
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

Short summary

This study decomposes which components of a coding agent—executable world models, simplification, and verification—drive performance on ARC-AGI-3. Testing four nested Codex-based agents across multiple GPT-5 variants, it finds that stronger models and greater reasoning effort consistently help, while individual component effects vary by setting. The full verification treatment ranks first everywhere and with gpt-5.6-sol solves every public game, though this likely reflects saturation of the public set only.

  • Stronger models and higher reasoning effort improve all agent variants on ARC-AGI-3
  • Verification treatment ranks first across all settings but uses substantially more resources
  • gpt-5.6-sol verification variant solves all public games, suggesting public-set saturation

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more