Vision Pretraining for Dense Spatial Perception

Published
Source
arXiv
Paper number
565
Field
Computer Vision
arXiv ID
2607.05247

Key points

  • It reframes boundaries as dense training signals rather than outputs and introduces boundary modeling into self-supervised pretraining.
  • Masked boundary modeling uses the teacher's boundary prediction as the masked target, and geometric routing resolves conflicts between semantics and boundaries.
  • A categorical reparameterization of the boundary field enables stable dense self-distillation.
  • A 1B LingBot-Vision model outperforms 7B DINOv3 on NYU-Depth v2 accuracy and reaches state-of-the-art on 14 depth-completion benchmarks.
  • A 0.3B distilled model matches the NYU-Depth v2 accuracy of 7B DINOv3 while using about 23 times fewer parameters.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)