AR
arXiv CS.AI
6/29/2026

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
Short summary
Researchers propose a three-stage training paradigm that teaches LLM agents to build internal world models for planning—overcoming their inherent reactivity in long-horizon tasks. The method combines latent capability injection, format-eliciting fine-tuning, and foresight-conditioned RL to produce grounded predictions. On search and mathematical reasoning tasks, the approach consistently outperforms standard agent training baselines.
- •Three-stage training pipeline enables LLM agents to internalize world models for proactive long-horizon planning
- •Combines World Model Agentic Mid-Training, Format-Eliciting SFT, and Foresight-Conditioned RL for grounded foresight
- •Demonstrates consistent improvements over baselines on search and mathematical reasoning tasks
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
