NTH

Training-free Behavior Cloning

AuthorsMaximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

October 3, 2026 2 min read
Watch on YouTube
The one-line take

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Key results

1000
MimicGen demonstrations

Training demonstrations used per tabletop task.

73%
BPC MimicGen mean success

Mean success rate across eight simulated tabletop tasks.

34
Real Switch successes

BPC successes out of 40 hardware trials, versus 8 for π0.5.

8
π0.5 Switch successes

π0.5 successes out of 40 hardware trials.

75 Hz
Jetson Orin Nano control rate

Onboard closed-loop drone control frequency.

20 hours
π0.5 hardware training time

Training time reported for real xArm6 manipulation.

What the paper found

Training-free Behavior Cloning introduces Behavior Predictive Control, or BPC, a robot imitation method that avoids end-to-end policy training by retaining raw demonstration trajectories and combining three components at runtime: action-aware retrieval, regularized Hankel-based history reconstruction, and a closed-form residual correction based on Random Fourier Features. Unlike Diffusion Policy, OpenVLA, or the vision-language-action model π0.5, BPC keeps every prediction traceable to supporting demonstrations, allowing operators to edit the demonstration bank, estimate task progress, and deploy across different robots and modalities, including RGB, proprioception, 3D inputs, and raw pixels. On MimicGen tabletop simulations using 1000 training demonstrations, BPC achieved a 73 percent mean success rate, competitive with π0.5 at 77 percent, while fitting in 1–5 minutes rather than hours. On real xArm6 manipulation, fitting took 30–120 seconds versus 20 hours for π0.5; notably, BPC solved the memory-dependent Switch task in 34 successes out of 40, compared with 8 out of 40 for π0.5. The method also controlled a drone onboard a Jetson Orin Nano at 75 Hz, demonstrating deployment on resource-constrained hardware. Its limitations are demonstration coverage, storage and retrieval costs, and vulnerability to ambiguous or out-of-distribution histories, but the central result is a modular behavior-cloning policy that links inference directly to editable data instead of opaque neural weights.

Original abstract

Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $π_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis