Back to feed
AR
arXiv CS.AI
7/21/2026
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

Short summary

Researchers propose masked diffusion language models (MDLMs) as text-based world models for agentic reinforcement learning, overcoming left-to-right bias of autoregressive models. MDLMs achieve better coherence and diversity than LLMs 4x their size, with up to 47% absolute gains on out-of-distribution environments using a GRPO training framework. The work is open-sourced with 239K trajectories across nine environments.

  • MDLMs outperform AR LMs as world models for agentic RL via bidirectional anchor-aware denoising
  • Plug-and-play GRPO framework achieves up to 47% gains on OOD environments without fine-tuning
  • 239K trajectories across 9 open-source environments and 12 model families, fully open-sourced

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more