NTH

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

AuthorsHaoyu Zhao, Xingyue Zhao, Hangyu Li, Biao Gong, Kehan Li, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li

July 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns a world model into a virtual teleoperation engine, letting humans generate robot training demonstrations without being physically connected to the robot.

Key results

40.0
Causal inference speed

Frames per second on a single NVIDIA H100 GPU.

1800
Robotic adaptation demonstrations

Real-world bimanual demonstration episodes used for robotic domain adaptation.

550
RynnWorld-Teleop FVD

FVD achieved by the full model on the action-conditioned evaluation benchmark.

20%
Lid Placement improvement

Success-rate increase when π0.5 or π0 is augmented with 300 synthetic episodes.

82.86%
Zero-real-data Block Pushing success

Success rate of Physical Intelligence’s π0 trained on 300 generated episodes without real data.

What the paper found

RynnWorld-Teleop, developed by Alibaba Group’s DAMO Academy, the Hong Kong Embodied AI Lab, and CUHK, proposes digital teleoperation: instead of moving a physical robot, an operator’s 21-joint hand-pose stream conditions a robot-centric world model that generates egocentric execution video and synchronized robot actions from a single reference image. Built on Wan2.2-TI2V-5B, the system combines depth-aware skeletal rendering, progressive pretraining on human egocentric video followed by robotic adaptation, and streaming autoregressive distillation with causal flow matching and DMD. Its causal student generates 40.0 FPS on a single NVIDIA H100 GPU, using four denoising steps and a KV cache, enabling interactive control. Training includes 1800 real-world bimanual demonstration episodes, while large-scale human pretraining supplies manipulation priors. On the EgoDex benchmark, the full model reduces FVD to 550, compared with 1223 for direct supervised fine-tuning of the same base model. As a data engine, RynnWorld-Teleop improves downstream policies: augmenting π0.5 or π0 with 300 synthetic episodes raises Lid Placement success by 20%, and Physical Intelligence’s π0 trained with 300 generated episodes and no real data reaches 82.86% on Block Pushing. These results suggest that digital teleoperation can scale robot-learning data collection without robot hardware, 3D assets, or repeated physical environment resets, although per-robot fine-tuning and weak performance on deformable or liquid interactions remain limitations.

Original abstract

Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator's hand-pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. We instantiate this paradigm in RynnWorld-Teleop, a system that integrates depth-aware skeletal conditioning, progressive human-to-robot training on a video Diffusion Transformer, and streaming autoregressive distillation. This pipeline compresses the generative process into a single-pass inference, enabling 40+ FPS, real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleop-generated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks. Moreover, augmenting real-world datasets with our digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis