RoboTTT: Context Scaling for Robot Policies
AuthorsYunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan
Resources
RoboTTT gives robot foundation models an 8K-step memory by learning at test time, enabling stronger imitation, adaptation, and long-horizon manipulation.
Key results
Maximum visuomotor context scaled without increasing inference latency
RoboTTT score across the three long-horizon assembly tasks
Relative improvement over GR00T N1.7 with single-step context
Improvement from 1K to 8K pretraining context, with scores of 43.9% and 71.5%
Fully successful Circuit rollouts out of 10 using one in-context human video
RoboTTT improvement from using failures as context and human corrections as targets
What the paper found
RoboTTT, from NVIDIA researchers collaborating with Stanford University and The University of Texas at Austin, adds Test-Time Training to NVIDIA’s GR00T N1.7 robot foundation model, replacing a conventional recurrent state with fast weights updated by gradient descent during both training and deployment. Its two-layer MLP fast weights compress long visuomotor histories into parameters, while sequence action forcing and truncated backpropagation through time make training feasible at 8K timesteps without increasing inference latency. On YAM bimanual-robot assembly tasks, RoboTTT reaches a 79% average task-completion score versus 42% for the single-step GR00T N1.7 baseline, an 87% improvement, and it is the only evaluated method to complete the five-minute, ten-stage Gear Bot task, succeeding in 2 of 10 trials. Scaling pretraining context from 1K to 8K timesteps raises closed-loop performance from 43.9% to 71.5%, a 63% gain, showing context length as a new scaling axis. Long-context conditioning also enables one-shot imitation from a single human video, with 6 of 10 Circuit trials fully successful, while the recurrent GDN baseline achieves 0 of 10. Finally, DAgger Distillation uses failed robot actions as context and human corrections as targets, producing a 36% improvement for RoboTTT and enabling on-the-fly recovery from execution errors.
Original abstract
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.