arXiv cs.CL
7/20/2026

Process Reward Informed Tree Rollout for Effective Multi-Turn RL
Short summary
PATR (Process-Scorer Guided Adaptive Tree Rollout) reframes multi-turn agent RL exploration as a tree-branching problem, using process feedback to score partial trajectories and selectively branch from promising states. It reuses shared prefixes and conservatively stops degenerate paths to reduce wasted sampling under the same training budget. Experiments show up to +5.0 points on SWE-Bench and +9.3 points on FrozenLake, demonstrating process-guided tree rollouts as an effective strategy for scalable multi-turn RL.
- •PATR organizes multi-turn RL trajectories as trees, branching from promising states using process reward scores
- •Reduces wasted sampling by stopping degenerate paths and reusing shared prefixes under the same budget
- •Improves SWE-Bench by up to +5.0 points and FrozenLake by +9.3 points over standard rollout methods
Generated with AI, which can make mistakes.
Is this a good recommendation for you?