Back to feed
AR
arXiv CS.AI
6/30/2026
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

Short summary

BV-Blend improves critic-free reinforcement learning by combining prompt-local reward statistics with semantic-cluster-conditioned historical moments, weighted by confidence scores. This stabilizes advantage estimation when reward variance is zero, addressing cold-start failures in binary verifier regimes. Experiments demonstrate improved training stability on verifiable reasoning benchmarks.

  • Addresses GRPO's instability in critic-free RL by blending on-policy and historical reward statistics
  • Uses confidence-weighted baseline adjustment to handle zero-variance scenarios in cold-start learning
  • Shows improved performance on verifiable reasoning tasks with reduced compute overhead vs. critic-based methods

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more