Vision Pretraining for Dense Spatial Perception
- Published
- Source
- arXiv
- Paper number
- 565
- Field
- Computer Vision
- arXiv ID
- 2607.05247
Key points
- It reframes boundaries as dense training signals rather than outputs and introduces boundary modeling into self-supervised pretraining.
- Masked boundary modeling uses the teacher's boundary prediction as the masked target, and geometric routing resolves conflicts between semantics and boundaries.
- A categorical reparameterization of the boundary field enables stable dense self-distillation.
- A 1B LingBot-Vision model outperforms 7B DINOv3 on NYU-Depth v2 accuracy and reaches state-of-the-art on 14 depth-completion benchmarks.
- A 0.3B distilled model matches the NYU-Depth v2 accuracy of 7B DINOv3 while using about 23 times fewer parameters.
Paper links
External research summaries. These are not HDATF publications or measured product results.