MarkTechPost
7/8/2026

Ant Group’s Robbyant Open-Sources LingBot-Vision: A 1B Boundary-Centric Vision Foundation Model for Dense Spatial Perception
Short summary
Ant Group's Robbyant has open-sourced LingBot-Vision, a 1B-parameter self-supervised ViT family designed for dense spatial perception using masked boundary modeling. The model treats image boundaries as a native training signal and reportedly matches or exceeds larger models on relevant benchmarks. It also serves as the initialization for LingBot-Depth 2.0, a downstream depth-estimation model.
- •LingBot-Vision is a 1B self-supervised ViT for dense spatial perception
- •Uses masked boundary modeling to make image boundaries a native training signal
- •Backbone initializes LingBot-Depth 2.0 and matches or surpasses larger models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


