FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

Published
Source
arXiv
Paper number
765
Field
Computer Vision
arXiv ID
2607.27110

Key points

  • It analyzes autoregressive video-generation error accumulation in the frequency domain and shows that low-frequency energy drift is the root cause.
  • Spectral Self-Anchoring, or SSA, fuses low-frequency content from the first frame with high-frequency content from the most recent frame to preserve both stability and dynamics.
  • Without extra training, it can extend a 5-second training model to stable generation for up to 2 minutes, which is a 24x horizon expansion.
  • On the VBench-Long metric, it outperforms previous training-free methods and remains competitive with training-based methods.
  • It adds only about 16.5% inference-latency overhead over Self-Forcing.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)