NTH

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

AuthorsBaochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen

June 21, 2026 2 min read
Watch on YouTube
The one-line take

LabVLA teaches vision-language-action models to execute scientific lab protocols with robot arms, aiming to make AI capable of carrying out experiments, not just planning them.

Key results

2947
LabAssetLibrary size

Annotated 3D assets generated by RoboGenesis

1000
LabTextureLibrary size

Curated texture images used for scene material assignment

10000
Laboratory scenes generated

Validated scenes synthesized by RoboGenesis

16
RoboGenesis robot platforms

Supported robot embodiments for cross-embodiment deployment

71.1%
LabUtopia average success ID

LabVLA average success rate in-distribution

70.0%
LabUtopia average success OOD

LabVLA average success rate out-of-distribution

What the paper found

LabVLA grounds vision-language-action learning in scientific laboratories by pairing a Qwen3-VL-4B-Instruct backbone with a DiT action expert and a synthetic data engine called RoboGenesis, developed by Zhejiang University and Shanghai AI Laboratory. RoboGenesis generates 2,947 annotated assets, 1,000+ textures, and 10,000 validated laboratory scenes, then composes long-horizon workflows from atomic skills such as pick, pour, press, open, and stir across 16 robot platforms. The model is trained in two stages: FAST action-token pretraining aligns the VLM with action semantics, and flow-matching posttraining with knowledge insulation decouples language grounding from continuous control; inference uses only 10 Euler steps. On LabUtopia, LabVLA reaches 71.1% average success in-distribution and 70.0% out-of-distribution, outperforming the next best policy, π0, by 7.8 and 6.8 percentage points. The synthetic corpus is transferable: fine-tuning X-VLA on LabEmbodied-Data lifts its five-task average from 49.3% to 64.3% ID and from 43.7% to 63.0% OOD. A physical Franka study on four 2–4-step laboratory tasks, each evaluated over 50 rollouts per condition, shows that the simulation-trained policy transfers to real benchtop manipulation, while Pour Liquid remains the hardest task across both simulation and hardware.

Original abstract

Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks alongside model design. To address the data side, we build RoboGenesis, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles. On the policy side, we present LabVLA, trained with a two-stage recipe: FAST action token pretraining first makes the Qwen3-VL-4B-Instruct backbone action aware before any continuous control is learned, and flow matching posttraining then attaches a DiT action expert under knowledge insulation. On the LabUtopia benchmark, LabVLA achieves the highest average success rate among all evaluated baselines under both in-distribution and out-of-distribution settings.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis