Self-Supervised Learning of Structured Dynamics from Videos

Published
Source
arXiv
Paper number
712
Field
Computer Vision
arXiv ID
2607.21576

Key points

  • The method proposes SDM, which structures video motion into two tokens, one for camera motion as the primary token and one for object motion as the residual token.
  • A lightweight model is placed on top of a frozen DINOv2 image backbone, which lets it learn dynamics representations without video pretraining.
  • It creates a new benchmark called ProbeMotion and measures performance systematically on synthetic and real videos.
  • With only weak supervision, it reaches performance similar to or better than VGGT trained on 17 datasets.
  • The primary token adapts automatically so that it captures the camera in camera-motion scenes and objects in static-camera scenes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)