Back to feed
AR
arXiv CS.AI
6/29/2026
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

Short summary

Researchers propose a three-stage training paradigm that teaches LLM agents to build internal world models for planning—overcoming their inherent reactivity in long-horizon tasks. The method combines latent capability injection, format-eliciting fine-tuning, and foresight-conditioned RL to produce grounded predictions. On search and mathematical reasoning tasks, the approach consistently outperforms standard agent training baselines.

  • Three-stage training pipeline enables LLM agents to internalize world models for proactive long-horizon planning
  • Combines World Model Agentic Mid-Training, Format-Eliciting SFT, and Foresight-Conditioned RL for grounded foresight
  • Demonstrates consistent improvements over baselines on search and mathematical reasoning tasks

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more