Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
AuthorsFilippo Lazzati, Kyle Stachowicz, William Chen, Alberto Maria Metelli, Andrew Wagenmaker, Sergey Levine
Resources
Action chunking helps robots not just by planning farther ahead, but by implicitly ensembling different temporal views of the task.
Key results
Number of tasks in the primary Libero benchmark suite.
Successful human demonstrations provided for each Libero task.
Success rate of the single-step Markovian behavioral cloning policy.
Success rate achieved by the best delayed policy, exceeding standard action chunking.
Success rate of the explicit ensemble compared with 12.6% for action chunking.
What the paper found
In “Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control,” Filippo Lazzati, Kyle Stachowicz, William Chen, Alberto Maria Metelli, Andrew Wagenmaker, and Sergey Levine, from Politecnico di Milano and UC Berkeley, analyze why robotic policies improve when they predict action sequences rather than single actions. Using diffusion policies on the Libero and Robomimic benchmarks, plus real-world Franka Emika experiments, they find that three common explanations—temporal consistency, horizon reduction, and representation learning—do not fully account for the gains. Instead, action chunking provides non-Markovian expressivity, reduces compounding error by predicting from past observations, and creates implicit ensembling: each chunk learns multiple relationships between an action and observations at different delays. On Libero-90, which contains 90 tasks and 50 human demonstrations per task, a Markovian policy reaches 68.9 percent success, while action chunking reaches 89.2 percent and a delayed policy reaches 94.0 percent, showing that chunk execution is often unnecessary. On harder Robomimic tasks, randomized-delay deployment of an action-chunked policy recovers comparable performance, while explicitly training ensembles can exceed it; on Robomimic Transport, explicit ensembling raises success from 12.6 percent for action chunking to 41.5 percent, an approximately 30 percent boost. The central conclusion is that action chunking works less because actions must be executed open-loop and more because it combines delayed prediction with ensemble-like robustness.
Original abstract
Action chunking---predicting and executing multiple actions instead of a single action---has proven to be a critical component for learning effective robotic control policies. However, our precise understanding of why action chunking improves performance has remained limited. In this work we seek to close this gap. Through rigorous experimental evaluations in both simulated and real-world settings, we show that existing hypotheses for the success of action chunking---temporal consistency, horizon reduction, and representation learning---fail to explain the success of action chunking. Instead, we find that action chunking benefits from greater non-Markovian expressivity and reduced compounding error compared to Markovian policies, but, in many settings of interest, these effects can be fully captured by delayed policies, which at each step predict a single action based on the observation $k$ steps in the past. We then show that there exists an additional benefit of action chunking that we refer to as implicit ensembling. In particular, by learning a diversity of temporal relationships (that is, $a_t | o_t, a_t | o_{t-1}, \ldots$), action-chunked policies exhibit behavior matching that of a model ensemble, increasing their robustness and generalization ability over policies that only learn a single temporal relationship. Building on these insights, we show that in simulated and real-world robotic control settings, we can match the performance of action chunking without action chunking---by deploying an action chunking policy as an ensemble of policies with randomized delays. Furthermore, we propose a policy class that amplifies the benefits of action chunking by explicitly instantiating an ensemble, and which we show significantly improves over the performance of action chunking in many domains.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.