arXiv cs.LG
7/29/2026

Inverse RL Helps Align AI by Imitating Humans
Short summary
PARED recovers an implicit reward function from expert demonstrations using a lightweight discriminator, eliminating the need for task-specific preference annotations. The recovered reward improves base policies through inference-time reranking and adversarial on-policy RL, and can be optimized after standard SFT for further gains. PARED also supports contextual alignment, tailoring a single policy to different audience preferences.
- •PARED extracts implicit rewards from demonstrations without preference annotations via a feature-space discriminator
- •Recovered reward improves policies through reranking and on-policy RL, with gains after SFT
- •Supports contextual alignment—single policy adapted to different audience preferences
Generated with AI, which can make mistakes.
Is this a good recommendation for you?