Back to feed
AR
arXiv CS.AI
7/16/2026
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

Short summary

Researchers introduce DROPJ, a human-centred method for safely training and deploying agent policies when environment dynamics are unknown and no reward function exists. A learned world model simulates trajectories, humans provide preferences with justifications, and a reward model guides model predictive control deployment. Experiments show preferences outperform other feedback types and safety justifications significantly enhance deployment safety.

  • DROPJ learns safe agent policies from human preferences and justifications via a learned world model
  • Human-generated simulated trajectories reduce training cost and improve deployment performance
  • Safety justifications accompanying preferences significantly enhance safety during deployment

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more