DINOcular: Self-Supervised Visuospatial Representations
- Published
- Source
- arXiv
- Paper number
- 1041
- Field
- Computer Vision
- arXiv ID
- 2608.27226
Key points
- It introduced depth as the third axis of rotary positional encoding (RoPE), allowing geometric information to enter feature representations naturally.
- Through self-supervised learning, learning without labels, it learned representations capturing both appearance and spatial structure from RGB-D video.
- It outperformed existing models of the same size on 3D correspondence matching in NAVI and ScanNet, and even surpassed DINOv3 ViT-B on ScanNet.
- In 3D pose estimation, there were conditions under which the large model, DINOcular-L, outperformed every baseline.
- It identified a trade-off in which improved 3D understanding came at a slight cost to semantic performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.