AR
arXiv CS.AI
7/20/2026

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Short summary
This study tests whether hierarchical reviewer agents actually improve math problem-solving in multi-agent systems. Using 4,181 Omni-MATH problems with matched gpt-oss-120b actors, it finds that a precise reviewer in a planner-executor-reviewer pipeline does not guarantee better outcomes—critique uptake matters more than critique accuracy. Broadcast-style peer discussion outperforms structured review in harder problem tiers because critiques are more likely to change the next answer.
- •Reviewer precision and critique uptake are empirically separable in multi-agent math reasoning
- •Broadcast peer discussion beats planner-executor-reviewer pipeline on harder problems despite less precise reviewers
- •Embedding reviewer guidance in solver context partially improves follow-through but doesn't close the gap
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
