MarkTechPost
7/11/2026

Ant Group’s Robbyant Unveils LingBot-VA 2.0: A Causal Video-Action Model Built Natively for Physical AI
Short summary
Ant Group's Robbyant has released LingBot-VA 2.0, a video-action foundation model designed from scratch for physical embodiment rather than fine-tuned from a video generator. The model features Foresight Reasoning for predicting future states, re-grounds on real observations, and achieves 225 Hz asynchronous control using a causal DiT architecture and sparse-MoE video stream. The article promises a breakdown of the technical details but the body itself is a brief teaser with no substantive analysis.
- •LingBot-VA 2.0 is a Physical AI video-action foundation model built natively for embodiment
- •Key features include Foresight Reasoning, 225 Hz async control, causal DiT, and sparse-MoE video stream
- •Article body is a teaser with no in-depth technical breakdown despite promising one
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


