NTH

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

AuthorsKian Hosseinkhani, Qinhe Peng, George Shramko, Mehran Aghabozorgi, Jianing Qian, Tristan Engst, Alireza Moazeni, Dinesh Jayaraman, Ke Li

September 13, 2026 2 min read
Watch on YouTube
The one-line take

IMLE-VLA replaces slow iterative action sampling with a single-step multimodal generator, making vision-language-action robots faster, smoother, and more effective.

Key results

55
Inference frequency

IMLE-VLA reaches 55 Hz versus 15 Hz for π0.5.

3.67
Inference speedup

The single-step action head increases inference frequency by 3.67 times.

11.0
Maximum action throughput

A longer execution horizon raises action throughput to 11.0 times the π0.5 baseline.

98.0%
LIBERO average success

Average success rate across the 40-task LIBERO benchmark.

3.0
Maximum jerk reduction

Real-world motion jerk is reduced by up to 3.0 times relative to π0.5.

6.6
Maximum inference-time reduction

VLA-only episode inference time is reduced by up to 6.6 times on the real robot.

What the paper found

IMLE-VLA addresses the inference bottleneck in vision-language-action policies by replacing π0.5’s 10-step flow-matching action head with a single-step conditional generator trained using conditional Implicit Maximum Likelihood Estimation, or cIMLE. The frozen vision-language backbone supplies the observation representation, while Gaussian noise produces multiple candidate action chunks; nearest-neighbor assignment trains candidates to cover different valid behaviors instead of collapsing multimodal actions into a regression mean. This drop-in action-head replacement raises inference frequency from 15 Hz to 55 Hz, a 3.67 times increase, and reaches 11.0 times higher action throughput with a longer execution horizon. On the 40-task LIBERO benchmark, IMLE-VLA achieves a 98.0% average success rate, exceeding π0.5 and faster alternatives such as Shallow-π and OpenVLA-OFT. Unlike OpenVLA-OFT, which uses a 7B backbone, IMLE-VLA preserves robustness on LIBERO-plus with a smaller 3B backbone. Real-world tests on a Franka Emika Panda, using NVIDIA hardware and the DROID dataset, show lower motion jerk by up to 3.0 times and reduce VLA-only episode inference time by up to 6.6 times, while outperforming π0.5 on all four tasks. The result is a teacher-free, multimodal, single-pass action head that improves responsiveness without retraining the full VLA backbone.

Original abstract

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $π_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse of naive regression heads while eliminating multi-step sampling entirely. When IMLE-VLA is applied to $π_{0.5}$, it increases inference frequency 3.67x (55 Hz vs. 15 Hz), enabling up to 11x higher action throughput. On the 40-task LIBERO benchmark, IMLE-VLA achieves the highest average success rate (98.0%) among all baselines while leading in inference frequency. Under the test-time perturbations of LIBERO-plus, IMLE-VLA retains $π_{0.5}$'s robustness while other baselines degrade sharply, confirming that the cIMLE head preserves generalization. Real-world experiments on a Franka Emika Panda across four tasks demonstrate smoother motion (2.2x to 3.0x lower jerk) and faster task completion, with IMLE-VLA outperforming $π_{0.5}$ on every task and reducing average VLA inference time per episode by 3.9x to 6.6x. Videos and code are available at https://kianhk6.github.io/IMLE-VLA/

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis