LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
AuthorsBaochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen
Resources
LabVLA teaches vision-language-action models to execute scientific lab protocols with robot arms, aiming to make AI capable of carrying out experiments, not just planning them.
Key results
Annotated 3D assets generated by RoboGenesis
Curated texture images used for scene material assignment
Validated scenes synthesized by RoboGenesis
Supported robot embodiments for cross-embodiment deployment
LabVLA average success rate in-distribution
LabVLA average success rate out-of-distribution
What the paper found
LabVLA grounds vision-language-action learning in scientific laboratories by pairing a Qwen3-VL-4B-Instruct backbone with a DiT action expert and a synthetic data engine called RoboGenesis, developed by Zhejiang University and Shanghai AI Laboratory. RoboGenesis generates 2,947 annotated assets, 1,000+ textures, and 10,000 validated laboratory scenes, then composes long-horizon workflows from atomic skills such as pick, pour, press, open, and stir across 16 robot platforms. The model is trained in two stages: FAST action-token pretraining aligns the VLM with action semantics, and flow-matching posttraining with knowledge insulation decouples language grounding from continuous control; inference uses only 10 Euler steps. On LabUtopia, LabVLA reaches 71.1% average success in-distribution and 70.0% out-of-distribution, outperforming the next best policy, π0, by 7.8 and 6.8 percentage points. The synthetic corpus is transferable: fine-tuning X-VLA on LabEmbodied-Data lifts its five-task average from 49.3% to 64.3% ID and from 43.7% to 63.0% OOD. A physical Franka study on four 2–4-step laboratory tasks, each evaluated over 50 rollouts per condition, shows that the simulation-trained policy transfers to real benchtop manipulation, while Pour Liquid remains the hardest task across both simulation and hardware.
Original abstract
Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks alongside model design. To address the data side, we build RoboGenesis, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles. On the policy side, we present LabVLA, trained with a two-stage recipe: FAST action token pretraining first makes the Qwen3-VL-4B-Instruct backbone action aware before any continuous control is learned, and flow matching posttraining then attaches a DiT action expert under knowledge insulation. On the LabUtopia benchmark, LabVLA achieves the highest average success rate among all evaluated baselines under both in-distribution and out-of-distribution settings.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.