
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
Short summary
Researchers introduce two cost-effective agent architectures—the Explorer-Definer Pipeline and the Reflective Orchestrator—that lift DeepSeek V3.2's one-shot ARC-AGI-1 baseline from 15.50% to 67.25% pass@2 at $0.62/task without benchmark-specific fine-tuning or heavy test-time compute. The orchestrator autonomously re-explores transformations when hypotheses fail, and unbiased pass@k analysis confirms the system is generation-bound rather than selection-bound, meaning future gains require broader candidate generation rather than better ranking. An ablation also identifies the pipeline's think tool as a significant contributor, with its removal dropping pass@2 by 5.75 points.
- •Explorer-Definer Pipeline separates pattern discovery from program synthesis, reaching 57.50% pass@2 at $0.25/task on ARC-AGI-1
- •Reflective Orchestrator adds adaptive re-exploration, reaching 67.25% pass@2 at $0.62/task — a ~52-point lift over the 15.50% one-shot baseline
- •Unbiased pass@k analysis shows the system is generation-bound, not selection-bound, directing future improvement toward broader generation rather than better ranking
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
