Back to feed
arXiv cs.CL
arXiv cs.CL
7/20/2026
Process Reward Informed Tree Rollout for Effective Multi-Turn RL

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

Short summary

PATR (Process-Scorer Guided Adaptive Tree Rollout) reframes multi-turn agent RL exploration as a tree-branching problem, using process feedback to score partial trajectories and selectively branch from promising states. It reuses shared prefixes and conservatively stops degenerate paths to reduce wasted sampling under the same training budget. Experiments show up to +5.0 points on SWE-Bench and +9.3 points on FrozenLake, demonstrating process-guided tree rollouts as an effective strategy for scalable multi-turn RL.

  • PATR organizes multi-turn RL trajectories as trees, branching from promising states using process reward scores
  • Reduces wasted sampling by stopping degenerate paths and reusing shared prefixes under the same budget
  • Improves SWE-Bench by up to +5.0 points and FrozenLake by +9.3 points over standard rollout methods

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more