Back to feed
Dev.to
Dev.to
7/21/2026
LingBot-Map: Open-Source Streaming Monocular 3D Reconstruction at 10,000+ Frames

LingBot-Map: Open-Source Streaming Monocular 3D Reconstruction at 10,000+ Frames

Original: LingBot-Map runs 10,000 frames on monocular video — no

Short summary

Ant Group's embodied-AI unit open-sourced LingBot-Map, a streaming monocular 3D reconstruction model that processes RGB video frame-by-frame at ~20 FPS using a paged KV-cache via FlashInfer. Its Geometric Context Attention maintains three memory pools (anchor, local pose-reference, compressed trajectory recap) to cut per-frame context growth ~80x versus full causal attention, enabling 10,000-frame inference without KV blowup. Benchmarks show 6.42m ATE on Oxford Spires, outperforming DA3 and VGGT. Three checkpoints are available on Hugging Face under Apache 2.0.

  • LingBot-Map: open-source streaming monocular 3D reconstruction at ~20 FPS, 10,000+ frames without KV cache explosion
  • Geometric Context Attention uses three memory pools to cut per-frame context growth ~80x vs full causal attention
  • Apache 2.0 release with three checkpoints on Hugging Face; benchmarks show 6.42m ATE on Oxford Spires, beating DA3 (12.87m) and VGGT (24.78m)

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more