Back to feed
arXiv cs.LG
arXiv cs.LG
8/4/2026
Inference-Time Policy Alignment for Fair Reinforcement Learning

Inference-Time Policy Alignment for Fair Reinforcement Learning

Short summary

This paper proposes inference-time policy shaping for aligning pretrained RL agents toward welfare-based fairness objectives without retraining. A multiplicative framework adjusts action probabilities using action-dependent welfare scores, compatible with any deep RL agent. Experiments across multiple domains show substantial fairness improvements while preserving core task performance.

  • Multiplicative policy shaping adjusts RL action probabilities at inference time using welfare scores—no retraining needed
  • Framework is compatible with any deep RL agent and preserves core task performance
  • Inspired by inference-time alignment techniques from large language models

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more