Learning High-Frequency Continuous Action Chunks in Latent Space
AuthorsKunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao, Wenchao Ding
The paper makes robot control smoother and more continuous by learning fast action chunks in latent space and refining them on the fly for real-world contact tasks.
Key results
High-frequency control setting used throughout the paper
Prediction horizon for each action chunk
Diagonal-Gaussian VAE latent size
Original action-space policy on synchronous execution
Latent-space counterpart on synchronous execution
What the paper found
Learning High-Frequency Continuous Action Chunks in Latent Space argues that robotic action chunking breaks down at 60 Hz when policies predict directly in action space, producing jitter, quantization error, and boundary stalls. The paper’s core idea is to encode 48-step high-frequency action chunks with a variational autoencoder, train the policy in a 10-dimensional latent space, and decode back to actions at execution time; this consistently improves trajectory smoothness and precision across Diffusion Policy, OpenVLA-OFT, and PI0.5. On real-world contact-rich tasks—Peel Cucumber, Wipe Vase, and Write Board—the latent formulation reduces jerk sharply; for example, OpenVLA-OFT on Write Board drops from 5.238 to 0.558 jerk and lifts success from 74% to 100%, while Diffusion Policy on the same task drops from 1.140 to 0.511 jerk. To handle asynchronous inference, the authors add Reuse-then-Refine, a training-free method that reuses overlapping executed actions, concatenates them with the new chunk, and refines the result through the VAE; this lowers chunk-boundary gaps and reduces execution stalls, especially for PI0.5 where the asynchronous write-board jerk falls from 4.984 to 1.754. The experiments also show that moderate latent compression helps, but overly aggressive compression degrades both precision and smoothness, and the VAE adds only about 2.30 ms of encode-decode overhead.
Original abstract
Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action frequencies, it becomes insufficient when the action frequency is further increased (e.g., to 60~Hz). At such high frequencies, policies often fail to generate actions that are both temporally smooth and spatially consistent. We address this challenge by shifting high-frequency action learning from the action space to a latent space with variational autoencoder (VAE). This formulation significantly improves both temporal and spatial consistency of high-frequency control. To enable smooth real-time execution, we further introduce Reuse-then-Refine, a chunk-level refine strategy that improves continuity between adjacent action chunks under asynchronous inference. As a result, robots controlled by our policy can execute complex contact-rich tasks continuously, with less pauses and jerky motions. Experiments on three real-world contact-rich robotic tasks show that our approach consistently completes tasks with smooth motions. Our code and data are available at https://github.com/tars-robotics/RTR.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.