Back to feed
MarkTechPost
MarkTechPost
7/9/2026
Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

Short summary

Ant Group's Robbyant released LingBot-VLA 2.0, an Apache-2.0 6B vision-language-action model for cross-embodiment robot manipulation. Pretrained on ~60,000 hours of data across 20 robot configurations, it maps all embodiments into a single 55-dimensional canonical action space using a token-level MoE action expert. It outperforms π0.5 and LingBot-VLA-1.0 on the GM-100 generalist benchmark.

  • Open-source 6B VLA model for cross-embodiment robot manipulation
  • Pretrained on 60K hours of robot trajectories and egocentric human video
  • Outperforms π0.5 on GM-100 benchmark across both evaluated platforms

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more