Back to feed
arXiv cs.CL
arXiv cs.CL
7/10/2026
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

Short summary

Reinforcement learning for LLMs often assigns equal positive credit to all tokens, including rare erroneous ones—a flaw called Positive-Credit Contamination. Researchers propose TACO, which calibrates credit by assessing each token's reliability risk, dampening noise while preserving useful rare patterns. Across three LLMs and eight benchmarks, TACO outperforms GRPO-style methods and significantly improves training stability.

  • Identifies Positive-Credit Contamination: low-probability tail tokens receive identical reinforcement as plausible ones, causing flawed reasoning
  • Proposes TACO: computes tail-risk scores to distinguish contextual errors from exploration, tuning positive credit without removing gradients
  • Results: outperforms GRPO baselines across 3 LLMs and 8 benchmarks with better stability and sustained long-horizon performance

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more