DINOcular: Self-Supervised Visuospatial Representations

Published
Source
arXiv
Paper number
1041
Field
Computer Vision
arXiv ID
2608.27226

Key points

  • It introduced depth as the third axis of rotary positional encoding (RoPE), allowing geometric information to enter feature representations naturally.
  • Through self-supervised learning, learning without labels, it learned representations capturing both appearance and spatial structure from RGB-D video.
  • It outperformed existing models of the same size on 3D correspondence matching in NAVI and ScanNet, and even surpassed DINOv3 ViT-B on ScanNet.
  • In 3D pose estimation, there were conditions under which the large model, DINOcular-L, outperformed every baseline.
  • It identified a trade-off in which improved 3D understanding came at a slight cost to semantic performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)