Alignment Forum
7/18/2026

Endogenous Alignment
Short summary
The article draws an analogy between human moral development and AI alignment, arguing that humans progress from exogenous alignment (rewards/punishments) to endogenous alignment (internalized guilt, shame, fear) and ultimately to harmony with values. Current AI alignment methods like RLHF and SFT resemble exogenous alignment but lack the endogenous self-correction mechanisms humans possess. The author questions whether AI could follow a similar developmental trajectory from external control to internalized value alignment.
- •Human alignment progresses from exogenous (rewards/punishments) to endogenous (internalized emotions like guilt and shame) to harmony with values
- •Current AI alignment via RLHF/SFT is exogenous — weights encode training but there is no internal motivation to stay aligned
- •AI faces an alignment ceiling because it lacks the evolutionary instincts that make humans alignable
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

