Self-Supervised Learning of Structured Dynamics from Videos
- Published
- Source
- arXiv
- Paper number
- 712
- Field
- Computer Vision
- arXiv ID
- 2607.21576
Key points
- The method proposes SDM, which structures video motion into two tokens, one for camera motion as the primary token and one for object motion as the residual token.
- A lightweight model is placed on top of a frozen DINOv2 image backbone, which lets it learn dynamics representations without video pretraining.
- It creates a new benchmark called ProbeMotion and measures performance systematically on synthetic and real videos.
- With only weak supervision, it reaches performance similar to or better than VGGT trained on 17 datasets.
- The primary token adapts automatically so that it captures the camera in camera-motion scenes and objects in static-camera scenes.
Paper links
External research summaries. These are not HDATF publications or measured product results.