NTH

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

AuthorsArnav Kumar Jain, Yilin Wu, Jesse Farebrother, Gokul Swamy, Andrea Bajcsy

July 3, 2026 2 min read
Watch on YouTube
The one-line take

WEAVER is a faster, longer-horizon world model for robot manipulation that can help robots predict outcomes, improve policies, and plan actions more effectively in the real world.

Key results

928M
model size

WEAVER total parameters

250
fine-tuning trajectories

real trajectories collected across five tasks

0.870
evaluation correlation

Pearson correlation between imagined and real success rates after fine-tuning

38%
policy improvement

real-world success rate gain over π0.5 without additional real interaction

5-10x
planning speedup

test-time planning faster than Ctrl-World

What the paper found

WEAVER, from Arnav Kumar Jain, Yilin Wu, Jesse Farebrother, Gokul Swamy, and Andrea Bajcsy across Mila and Carnegie Mellon University, is a multi-view world model for robotic manipulation that tries to solve fidelity, long-horizon consistency, and efficiency at once. It encodes wrist and external camera streams plus proprioception with a pretrained Stable Diffusion 3 VAE, predicts future latents with a 32-layer, 1536-dimensional transformer trained using flow matching and diffusion forcing, and adds latent reward and critic heads distilled from Robometer to score imagined rollouts without an external VLM judge. Trained at 928M parameters on DROID for 1M steps, then fine-tuned on 250 real trajectories from five hardware tasks, WEAVER outperforms Ctrl-World on both in-distribution and OOD manipulation rollouts, including lower DROID exterior FID of 10.20 versus 26.09 at 16 NFE and lower DROID wrist FVD of 90.72 versus 195.37. On real robot tasks, the fine-tuned model reaches ρ = 0.870 correlation between imagined and real success rates for policy evaluation, improves the π0.5 policy’s real-world success rate by 38% without additional real interaction, and enables test-time planning with 5–10× faster inference than prior world models, reaching about 20× lower dynamics latency than Ctrl-World in the full planning pipeline. The paper’s main contribution is not just better video prediction, but a world-model stack that is actually usable for evaluation, policy steering, and online planning on contact-rich manipulation.

Original abstract

The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world interaction. To unlock these downstream capabilities, a WM needs to jointly satisfy three desiderata: $\textit{(i)}$ fidelity (i.e., producing simulated trajectories that correlate with reality), $\textit{(ii)}$ consistency (i.e., producing simulated trajectories that are coherent over long horizons), and $\textit{(iii)}$ efficiency (i.e., producing simulated trajectories quickly). We propose $\texttt{WEAVER}$ (World Estimation Across Views for Embodied Reasoning): a WM architecture that simultaneously achieves all three desiderata, providing state-of-the-art results on robotic manipulation tasks. $\texttt{WEAVER}$ is a multi-view WM trained to predict future latents and reward values via a flow-matching loss. We distill the key design decisions across model architecture, memory, and prediction objectives required to unlock the kinds of long-horizon dynamic manipulation tasks that have confounded prior world modeling approaches. We apply $\texttt{WEAVER}$ in robotic hardware, demonstrating its effectiveness at policy evaluation ($ρ$=0.870 correlation with real-world success rate), policy improvement (real-world success rate improvement of $38\%$ on top of the $π_{0.5}$ robot foundation model), and test-time planning (real-world success rate improvement of $14\%$ with a $5-10\times$ speedup over prior WMs). $\texttt{WEAVER}$ also demonstrates better performance than prior WMs when evaluated on out-of-distribution scenarios. Code, models, and videos at: https://arnavkj1995.github.io/WEAVER/ .

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis