FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
- Published
- Source
- arXiv
- Paper number
- 765
- Field
- Computer Vision
- arXiv ID
- 2607.27110
Key points
- It analyzes autoregressive video-generation error accumulation in the frequency domain and shows that low-frequency energy drift is the root cause.
- Spectral Self-Anchoring, or SSA, fuses low-frequency content from the first frame with high-frequency content from the most recent frame to preserve both stability and dynamics.
- Without extra training, it can extend a 5-second training model to stable generation for up to 2 minutes, which is a 24x horizon expansion.
- On the VBench-Long metric, it outperforms previous training-free methods and remains competitive with training-based methods.
- It adds only about 16.5% inference-latency overhead over Self-Forcing.
Paper links
External research summaries. These are not HDATF publications or measured product results.