You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences

Published
Source
arXiv
Paper number
432
Field
Computer Vision
arXiv ID
2606.15956

Key points

  • It is the first approach to self-supervised learning that does so with only a causality assumption and without strong inductive biases such as augmentation, masking, or cropping.
  • It trains the frame encoder and motion encoder jointly so that current frame plus motion delta equals the next frame representation.
  • It experimentally verifies the hypothesis that weaker inductive bias is better as data scale increases.
  • It achieves performance comparable to DINO and iBOT on ADE-20K semantic segmentation and MPI-Sintel optical flow.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)