Self Gradient Forcing: Native Long Video Extrapolation

Published
Source
arXiv
Paper number
690
Field
Computer Vision
arXiv ID
2607.20368

Key points

  • It defines and resolves the past-context gradient gap, which is the critical weakness of Self Forcing.
  • A two-pass strategy, where the first pass records rollouts and the second pass reconstructs gradients in parallel, solves the problem without extra memory overhead.
  • Using only 5-second training videos, it can generate coherent long videos up to 240 seconds, improving subject identity, background consistency, and temporal stability.
  • A user study found that it was preferred over Self Forcing in every setting, with GSB ranging from 29.6% to 48.7%.
  • Peak memory increases by only 8 GB, and training slowdown is about 13%, which makes the method practical.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)