Alignment Forum
7/12/2026
From wantons to moral agents
Short summary
A theoretical post exploring how AI agents might transition from wanton behavior (driven by first-order desires like fixed RL rewards) to moral agents capable of reflective endorsement. Drawing on Frankfurt's 1971 hierarchy of desires, the author examines mechanisms by which agents could develop second-order preferences about their own reasoning processes. The analysis connects philosophical frameworks to concrete AI examples like RL agents with modifiable reward functions.
- •Applies Frankfurt's 'wanton' concept to AI agents with fixed RL rewards
- •Explores mechanisms for transition from first-order desire-driven behavior to reflective endorsement
- •Connects philosophical frameworks to concrete AI agent architectures
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

