Back to feed
AR
arXiv CS.AI
7/21/2026
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

Short summary

W2SPO is an off-policy RL method that injects short auxiliary segments (as brief as 8 tokens) from a weaker model into a target model's reasoning trajectory to escape semantic redundancy bottlenecks. Policy updates are restricted to inserted segments based on final verifiable rewards. It improves Pass@1 from 62.3% to 64.2% over vanilla GRPO with a 3.55x training speedup on 4B-scale math reasoning benchmarks.

  • W2SPO injects 8-token auxiliary segments from weaker models to diversify target model exploration
  • Improves Pass@1 from 62.3% to 64.2% over GRPO with 3.55x speedup on math benchmarks
  • Addresses the support-limited bottleneck where self-generated rollouts converge to erroneous reasoning basins

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more