You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- Published
- Source
- arXiv
- Paper number
- 432
- Field
- Computer Vision
- arXiv ID
- 2606.15956
Key points
- It is the first approach to self-supervised learning that does so with only a causality assumption and without strong inductive biases such as augmentation, masking, or cropping.
- It trains the frame encoder and motion encoder jointly so that current frame plus motion delta equals the next frame representation.
- It experimentally verifies the hypothesis that weaker inductive bias is better as data scale increases.
- It achieves performance comparable to DINO and iBOT on ADE-20K semantic segmentation and MPI-Sintel optical flow.
Paper links
External research summaries. These are not HDATF publications or measured product results.