AR
arXiv CS.AI
6/30/2026

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
Short summary
BV-Blend improves critic-free reinforcement learning by combining prompt-local reward statistics with semantic-cluster-conditioned historical moments, weighted by confidence scores. This stabilizes advantage estimation when reward variance is zero, addressing cold-start failures in binary verifier regimes. Experiments demonstrate improved training stability on verifiable reasoning benchmarks.
- •Addresses GRPO's instability in critic-free RL by blending on-policy and historical reward statistics
- •Uses confidence-weighted baseline adjustment to handle zero-variance scenarios in cold-start learning
- •Shows improved performance on verifiable reasoning tasks with reduced compute overhead vs. critic-based methods
Generated with AI, which can make mistakes.
Is this a good recommendation for you?