NTH

RoboTTT: Context Scaling for Robot Policies

AuthorsYunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan

July 18, 2026 2 min read
Watch on YouTube
The one-line take

RoboTTT gives robot foundation models an 8K-step memory by learning at test time, enabling stronger imitation, adaptation, and long-horizon manipulation.

Key results

8K
RoboTTT context length

Maximum visuomotor context scaled without increasing inference latency

79%
Average task completion

RoboTTT score across the three long-horizon assembly tasks

87%
Improvement over single-step baseline

Relative improvement over GR00T N1.7 with single-step context

63%
Closed-loop gain from context scaling

Improvement from 1K to 8K pretraining context, with scores of 43.9% and 71.5%

6
One-shot video imitation

Fully successful Circuit rollouts out of 10 using one in-context human video

36%
DAgger Distillation improvement

RoboTTT improvement from using failures as context and human corrections as targets

What the paper found

RoboTTT, from NVIDIA researchers collaborating with Stanford University and The University of Texas at Austin, adds Test-Time Training to NVIDIA’s GR00T N1.7 robot foundation model, replacing a conventional recurrent state with fast weights updated by gradient descent during both training and deployment. Its two-layer MLP fast weights compress long visuomotor histories into parameters, while sequence action forcing and truncated backpropagation through time make training feasible at 8K timesteps without increasing inference latency. On YAM bimanual-robot assembly tasks, RoboTTT reaches a 79% average task-completion score versus 42% for the single-step GR00T N1.7 baseline, an 87% improvement, and it is the only evaluated method to complete the five-minute, ten-stage Gear Bot task, succeeding in 2 of 10 trials. Scaling pretraining context from 1K to 8K timesteps raises closed-loop performance from 43.9% to 71.5%, a 63% gain, showing context length as a new scaling axis. Long-context conditioning also enables one-shot imitation from a single human video, with 6 of 10 Circuit trials fully successful, while the recurrent GDN baseline achieves 0 of 10. Finally, DAgger Distillation uses failed robot actions as context and human corrections as targets, producing a 36% improvement for RoboTTT and enabling on-the-fly recovery from execution errors.

Original abstract

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis