NTH

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

AuthorsYuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao

August 21, 2026 3 min read
Watch on YouTube
The one-line take

Evoke builds an interactive world model that remembers evolving scenes indefinitely while generating responsive video in just three steps.

Key results

30
Long-horizon supervision

Seconds covered by the distribution-matching teacher objective.

2.11
Chunk latency

Seconds to generate each 1.5-second chunk on one NVIDIA H200 at 384 × 640.

80.8
WBench overall

Overall score on the 158-case WBench navigation split.

66.77
VBench-2.0

Overall video-generation score using three sampling steps.

85.11
VBench-Long

Long-video benchmark score using three sampling steps.

65.5
Long-session evaluation

Duration of the quantitative continuous rollouts.

What the paper found

Alaya-EVOKE, or Evoke, is a three-step interactive video world model designed to generate indefinitely without allowing denoiser context to grow with session length. It externalizes persistent scene geometry in a camera-indexed world state bank: generated frames are lifted into geometry, stored under camera pose, and rendered back into later views, while bounded local history keeps each inference call fixed-cost. Its teacher is built on the 14B Wan2.2 A14B diffusion transformer and uses chunk-wise sparse attention, distant-frame retrieval, and a linear-attention global state, changing long-sequence computation from quadratic to approximately linear growth. A 30-second distribution-matching objective with self-forced rollouts transfers long-horizon stability and per-chunk prompt control to a classifier-free-guidance-free student using only three denoising evaluations. On the 158-case WBench navigation split, Evoke reaches 80.8 overall, ranking first in the reported leaderboard, while scoring 66.77 on VBench-2.0 and 85.11 on VBench-Long, competitive with systems such as Veo 3 and Sora despite their different sampling configurations. Across 65.5-minute rollouts, the system avoids progressively worsening photometric drift; on a single NVIDIA H200 at 384 × 640, diffusion generates each 1.5-second chunk in 2.11 seconds. The main limitation is that geometric memory preserves coarse revisited structure better than fine object identity, appearance, or dynamic state.

Original abstract

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at $384\times 640$, each $1.5\,\mathrm{s}$ chunk is generated in $2.11\,\mathrm{s}$. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis