MarkTechPost
7/9/2026

Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation
Short summary
Ant Group's Robbyant released LingBot-VLA 2.0, an Apache-2.0 6B vision-language-action model for cross-embodiment robot manipulation. Pretrained on ~60,000 hours of data across 20 robot configurations, it maps all embodiments into a single 55-dimensional canonical action space using a token-level MoE action expert. It outperforms π0.5 and LingBot-VLA-1.0 on the GM-100 generalist benchmark.
- •Open-source 6B VLA model for cross-embodiment robot manipulation
- •Pretrained on 60K hours of robot trajectories and egocentric human video
- •Outperforms π0.5 on GM-100 benchmark across both evaluated platforms
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


