arXiv cs.LG
7/30/2026

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
Short summary
MeRLa introduces a meta-learned reward shaping framework for RLHF that learns task-aware shaping functions across auxiliary tasks before RLHF training. The composite reward preserves policy optimality while providing task-specific learning signals. Experiments on LLaMA-3-8B show consistent improvements over PPO, DPO, GRPO, and DAPO, achieving 90.8% length-controlled win rate on AlpacaEval 2.0 and 9.14 on MT-Bench with 41% less training instability.
- •Meta-learned reward shaping produces task-aware composite rewards for RLHF
- •Outperforms PPO, DPO, GRPO, and DAPO on LLaMA-3-8B across four benchmarks
- •Achieves 90.8% AlpacaEval 2.0 win rate with 41% less training instability
Generated with AI, which can make mistakes.
Is this a good recommendation for you?