arXiv cs.LG
7/17/2026

RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Short summary
RENEW introduces Dynamics Learning from Human Feedback (DLHF) to repair world model exploitation in offline RL by using human preferences over imagined rollouts. Naive DLHF is sample-inefficient, so RENEW uses epistemic uncertainty to focus finetuning where the model is most exploitable. Evaluated on Jumanji and classic control environments, RENEW improves sample efficiency, limits catastrophic forgetting, and reduces exploitation in pretrained world models.
- •DLHF uses human preferences over imagined rollouts to correct world model hallucinations
- •RENEW focuses finetuning on high-uncertainty regions for practical sample efficiency
- •Demonstrated on Jumanji and classic control tasks with reduced exploitation and forgetting
Generated with AI, which can make mistakes.
Is this a good recommendation for you?