arXiv cs.LG
8/4/2026

Inference-Time Policy Alignment for Fair Reinforcement Learning
Short summary
This paper proposes inference-time policy shaping for aligning pretrained RL agents toward welfare-based fairness objectives without retraining. A multiplicative framework adjusts action probabilities using action-dependent welfare scores, compatible with any deep RL agent. Experiments across multiple domains show substantial fairness improvements while preserving core task performance.
- •Multiplicative policy shaping adjusts RL action probabilities at inference time using welfare scores—no retraining needed
- •Framework is compatible with any deep RL agent and preserves core task performance
- •Inspired by inference-time alignment techniques from large language models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?