NTH

Learning High-Frequency Continuous Action Chunks in Latent Space

AuthorsKunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao, Wenchao Ding

June 26, 2026 2 min read
Watch on YouTube
The one-line take

The paper makes robot control smoother and more continuous by learning fast action chunks in latent space and refining them on the fly for real-world contact tasks.

Key results

60
action frequency

High-frequency control setting used throughout the paper

48
action chunk length

Prediction horizon for each action chunk

10
latent dimension

Diagonal-Gaussian VAE latent size

5.238
OpenVLA-OFT write-board jerk

Original action-space policy on synchronous execution

0.558
OpenVLA-OFT write-board latent jerk

Latent-space counterpart on synchronous execution

What the paper found

Learning High-Frequency Continuous Action Chunks in Latent Space argues that robotic action chunking breaks down at 60 Hz when policies predict directly in action space, producing jitter, quantization error, and boundary stalls. The paper’s core idea is to encode 48-step high-frequency action chunks with a variational autoencoder, train the policy in a 10-dimensional latent space, and decode back to actions at execution time; this consistently improves trajectory smoothness and precision across Diffusion Policy, OpenVLA-OFT, and PI0.5. On real-world contact-rich tasks—Peel Cucumber, Wipe Vase, and Write Board—the latent formulation reduces jerk sharply; for example, OpenVLA-OFT on Write Board drops from 5.238 to 0.558 jerk and lifts success from 74% to 100%, while Diffusion Policy on the same task drops from 1.140 to 0.511 jerk. To handle asynchronous inference, the authors add Reuse-then-Refine, a training-free method that reuses overlapping executed actions, concatenates them with the new chunk, and refines the result through the VAE; this lowers chunk-boundary gaps and reduces execution stalls, especially for PI0.5 where the asynchronous write-board jerk falls from 4.984 to 1.754. The experiments also show that moderate latent compression helps, but overly aggressive compression degrades both precision and smoothness, and the VAE adds only about 2.30 ms of encode-decode overhead.

Original abstract

Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action frequencies, it becomes insufficient when the action frequency is further increased (e.g., to 60~Hz). At such high frequencies, policies often fail to generate actions that are both temporally smooth and spatially consistent. We address this challenge by shifting high-frequency action learning from the action space to a latent space with variational autoencoder (VAE). This formulation significantly improves both temporal and spatial consistency of high-frequency control. To enable smooth real-time execution, we further introduce Reuse-then-Refine, a chunk-level refine strategy that improves continuity between adjacent action chunks under asynchronous inference. As a result, robots controlled by our policy can execute complex contact-rich tasks continuously, with less pauses and jerky motions. Experiments on three real-world contact-rich robotic tasks show that our approach consistently completes tasks with smooth motions. Our code and data are available at https://github.com/tars-robotics/RTR.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis