NTH

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

AuthorsQixin Hu, Shuai Yang, Wei Huang, Song Han, Yukang Chen

June 14, 2026 3 min read
Watch on YouTube
The one-line take

LongLive-RAG turns previously generated video latents into a searchable memory, helping autoregressive models keep long videos more consistent instead of drifting over time.

Key results

1024
embedding dimension

Retrieval embedding size used by the latent autoencoder

4.08
retrieval overhead per block

Milliseconds added per autoregressive block during a 120s rollout

490
retrieval overhead total

Total retrieval time in milliseconds for a 120s rollout

6
best retrieval budget

Top-K retrieval setting that gave the strongest consistency and imaging quality

4.70
auxiliary VLM score

LongLive-RAG score on the 30s Causal-Forcing auxiliary video quality evaluation

What the paper found

LongLive-RAG, developed by NVIDIA researchers with affiliations including USC and MIT, reframes long video generation as retrieval-augmented generation over self-generated latents instead of relying only on sliding-window attention. The method keeps the base autoregressive video diffusion backbone frozen and adds a lightweight retrieval path: a query embedding from the latest completed latent searches a historical latent bank, then the top-K matched context entries are injected back into attention alongside the local window. To make retrieval useful rather than redundant, the paper trains a 1024-dimensional latent retrieval encoder with a reconstruction loss, a Window Temporal Delta Loss that suppresses near-duplicate short-range matches, and a second-order smoothing term to stabilize the embedding trajectory. Across three backbones—Causal-Forcing, Self-Forcing, and LongLive—and 30s, 60s, and 120s rollouts, LongLive-RAG achieves the best average VBench-Long rank in every setting, with especially strong gains in subject consistency, background consistency, motion smoothness, and imaging quality. The authors also report that the retrieval overhead is only 4.08 ms per block, or 490 ms total for a 120s rollout, and that K = 6 is the best retrieval budget under a fixed attention budget. On 30s Causal-Forcing, the full method improves auxiliary VLM exposure quality to 4.70 and beats random retrieval, average pooling, and reconstruction-only embeddings, supporting the claim that content-addressable memory is more effective than fixed anchors or compressed history for long-horizon video fidelity.

Original abstract

Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly adopt sliding-window attention during generation. This creates an irreversible generation trajectory: once the active window accumulates appearance errors, subsequent generations can only condition on this degraded trajectory and drift further away. We address this limitation by formulating long video generation as a retrieval-augmented generation (RAG) problem. Rather than relying solely on the recent window, we treat previously generated latents as a dynamic, searchable history. We propose LongLive-RAG, a general retrieval framework for AR video generation. At each new block, LongLive-RAG uses a query embedding to retrieve relevant historical latents. This lightweight retrieval step adds only a small overhead relative to generation and lets the generator condition on non-local context instead of only the recent window. To make retrieval more discriminative, we introduce the Window Temporal Delta Loss that suppresses redundant local similarity and encourages embeddings to capture meaningful temporal changes. Together, these components help reduce error accumulation caused by sliding-window attention. Experiments across multiple AR backbones and generation lengths show improved long-video quality and the best average VBench-Long rank. To our knowledge, among open-ended AR long video generation methods, LongLive-RAG is the first to formulate self-generated latent history as content-addressable retrieval memory. Code is available at https://github.com/qixinhu11/LongLive-RAG.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis