NTH

Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

AuthorsDylan Miller, Martin Jagersand

August 10, 2026 2 min read
Watch on YouTube
The one-line take

Temporal Policy makes diffusion-like robot control faster by starting action generation from the robot’s recent history instead of random noise.

Key results

19.1 ms
Inference latency

Temporal Policy inference on a single NVIDIA RTX 4080.

10
Function evaluations

Number of function evaluations used for action generation.

17M
Model size

Compact Temporal Policy architecture used in simulation experiments.

1.21
Square transport cost

Temporal Policy transport cost on the Square proficient-human task.

95%
Real-world grasp success

Successful mug acquisition across 20 physical trials.

50%
Real-world overall success

Overall mug-hanging success across 20 physical trials.

What the paper found

Temporal Policy is a generative policy for robotic learning from demonstration that replaces the independent Gaussian-noise initialization used by Diffusion Policy and standard flow matching with the robot’s recent state history. Built on stochastic interpolants, it formulates action generation as point-to-distribution transport, coupling past configurations directly to future action chunks so the learned vector field is shorter and straighter. A 1D U-Net predicts the interpolant drift, while a ResNet-18 encodes visual observations; analytic score recovery also allows deterministic ODE or stochastic SDE sampling without retraining. On the Robomimic benchmark, Temporal Policy used 10 function evaluations and a compact 17M-parameter model, compared with 255M-parameter baselines, achieving 19.1 ms inference on a single NVIDIA RTX 4080 while matching the success rates of Diffusion Policy and CFM. On the Square proficient-human task, its transport cost was 1.21 versus 189.99 for Diffusion Policy, with a straightness ratio of 1.02 versus 20.61. Physical evaluation on a 7-DoF Barrett WAM used 150 demonstrations and produced a 95% grasp success rate and 50% overall success rate across 20 mug-hanging trials, showing that history-initialized transport can support low-latency closed-loop control without sacrificing multimodal action generation.

Original abstract

By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learning from Demonstration (LfD), but often suffer from high inference cost. This paper introduces Temporal Policy, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem. By initializing the generative flow at the robot's recent history, we explicitly couple past states to future action sequences. This data-dependent coupling reduces transport cost and produces straight vector fields. We validate Temporal Policy across visuomotor simulation benchmarks and on a physical Barrett WAM 2x 7DoF teleoperation platform. Our approach reduces transport costs by nearly an order of magnitude compared to noise-initialized baselines, achieving a 19.1 ms inference latency on a single NVIDIA RTX 4080. Crucially, these geometric and computational efficiencies are achieved while matching the success rates of state-of-the-art baselines. This simplified transport geometry bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control. The code is publicly available at https://github.com/dmiller12/TemporalPolicy.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis