Back to feed
Alignment Forum
Alignment Forum
7/18/2026
Endogenous Alignment

Endogenous Alignment

Short summary

The article draws an analogy between human moral development and AI alignment, arguing that humans progress from exogenous alignment (rewards/punishments) to endogenous alignment (internalized guilt, shame, fear) and ultimately to harmony with values. Current AI alignment methods like RLHF and SFT resemble exogenous alignment but lack the endogenous self-correction mechanisms humans possess. The author questions whether AI could follow a similar developmental trajectory from external control to internalized value alignment.

  • Human alignment progresses from exogenous (rewards/punishments) to endogenous (internalized emotions like guilt and shame) to harmony with values
  • Current AI alignment via RLHF/SFT is exogenous — weights encode training but there is no internal motivation to stay aligned
  • AI faces an alignment ceiling because it lacks the evolutionary instincts that make humans alignable

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more