arXiv cs.CL
7/10/2026

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning
Short summary
Reinforcement learning for LLMs often assigns equal positive credit to all tokens, including rare erroneous ones—a flaw called Positive-Credit Contamination. Researchers propose TACO, which calibrates credit by assessing each token's reliability risk, dampening noise while preserving useful rare patterns. Across three LLMs and eight benchmarks, TACO outperforms GRPO-style methods and significantly improves training stability.
- •Identifies Positive-Credit Contamination: low-probability tail tokens receive identical reinforcement as plausible ones, causing flawed reasoning
- •Proposes TACO: computes tail-risk scores to distinguish contextual errors from exploration, tuning positive credit without removing gradients
- •Results: outperforms GRPO baselines across 3 LLMs and 8 benchmarks with better stability and sustained long-horizon performance
Generated with AI, which can make mistakes.
Is this a good recommendation for you?