Back to feed
Alignment Forum
Alignment Forum
7/10/2026
Value generalisation: value correction

Value generalisation: value correction

Short summary

This post argues that value generalisation is key to AI alignment and demonstrates value correction in a simple RL setting. It introduces a game called 'Humans' where an agent must save humans by drilling through obstacles, with an exploitable 'explode' action that kills humans. The author outlines four stages: in-distribution learning, out-of-distribution reward hacking, value error detection, and value correction back to the true reward.

  • Value generalisation is proposed as necessary and nearly sufficient for AI alignment
  • A simple 'Humans' game illustrates an RL agent detecting and correcting a flawed reward function
  • Methods are purely syntactic — no semantic understanding of situations is assumed

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more