AR
arXiv CS.AI
7/21/2026

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Short summary
Researchers propose masked diffusion language models (MDLMs) as text-based world models for agentic reinforcement learning, overcoming left-to-right bias of autoregressive models. MDLMs achieve better coherence and diversity than LLMs 4x their size, with up to 47% absolute gains on out-of-distribution environments using a GRPO training framework. The work is open-sourced with 239K trajectories across nine environments.
- •MDLMs outperform AR LMs as world models for agentic RL via bidirectional anchor-aware denoising
- •Plug-and-play GRPO framework achieves up to 47% gains on OOD environments without fine-tuning
- •239K trajectories across 9 open-source environments and 12 model families, fully open-sourced
Generated with AI, which can make mistakes.
Is this a good recommendation for you?