
LingBot-Map: Open-Source Streaming Monocular 3D Reconstruction at 10,000+ Frames
Original: LingBot-Map runs 10,000 frames on monocular video — no
Short summary
Ant Group's embodied-AI unit open-sourced LingBot-Map, a streaming monocular 3D reconstruction model that processes RGB video frame-by-frame at ~20 FPS using a paged KV-cache via FlashInfer. Its Geometric Context Attention maintains three memory pools (anchor, local pose-reference, compressed trajectory recap) to cut per-frame context growth ~80x versus full causal attention, enabling 10,000-frame inference without KV blowup. Benchmarks show 6.42m ATE on Oxford Spires, outperforming DA3 and VGGT. Three checkpoints are available on Hugging Face under Apache 2.0.
- •LingBot-Map: open-source streaming monocular 3D reconstruction at ~20 FPS, 10,000+ frames without KV cache explosion
- •Geometric Context Attention uses three memory pools to cut per-frame context growth ~80x vs full causal attention
- •Apache 2.0 release with three checkpoints on Hugging Face; benchmarks show 6.42m ATE on Oxford Spires, beating DA3 (12.87m) and VGGT (24.78m)
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


