AR
arXiv CS.AI
7/21/2026

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Short summary
W2SPO is an off-policy RL method that injects short auxiliary segments (as brief as 8 tokens) from a weaker model into a target model's reasoning trajectory to escape semantic redundancy bottlenecks. Policy updates are restricted to inserted segments based on final verifiable rewards. It improves Pass@1 from 62.3% to 64.2% over vanilla GRPO with a 3.55x training speedup on 4B-scale math reasoning benchmarks.
- •W2SPO injects 8-token auxiliary segments from weaker models to diversify target model exploration
- •Improves Pass@1 from 62.3% to 64.2% over GRPO with 3.55x speedup on math benchmarks
- •Addresses the support-limited bottleneck where self-generated rollouts converge to erroneous reasoning basins
Generated with AI, which can make mistakes.
Is this a good recommendation for you?