G0.5: One Autoregressive Stream for Robot Reasoning and Action
AuthorsYicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao
Resources
G0.5 turns a vision-language model into a single-stream robot brain that explains what it is doing while directly generating the actions to do it.
Key results
Average success after DROID post-training on unseen environments and object instances
Overall benchmark success rate
Average success across clean and randomized evaluation
Task Success Score across 50 household tasks after four post-training epochs
Average success on R1-Lite and R1-Pro manipulation settings
What the paper found
G0.5 replaces the prevailing VLA design—where a VLM such as Qwen3.5 conditions a separate flow-matching expert—with one autoregressive transformer that reasons and acts in the same token stream. Initialized from Qwen3.5 2B, it uses a learned cross-embodiment ActionCodec with residual vector quantization, a 27-dimensional structured action space, active-degree-of-freedom tokenization, native chain-of-thought for subtasks, bounding boxes, 2D traces, and action hints, plus factorized visual memory for multi-second history. A single next-token cross-entropy objective trains language, reasoning, and motor commands, while exact action-token likelihoods make reinforcement learning with GRPO more direct than for flow matching. Pretraining combines 14 robot embodiments, robot demonstrations, VQA data, and automated annotations generated with Gemini 3, Doubao Seed 2.0 Pro, and SAM3. After post-training, G0.5 reaches 82.5% on DROID with held-out environments and object instances, 98.9% on LIBERO, 93.3% on RoboTwin 2.0, and 87.3% on SimplerEnv-Bridge. On the 2025 BEHAVIOR Challenge, using one generalist checkpoint across 50 household tasks, it achieves 31.4%, exceeding π0.5’s 26.3%; real-world fine-tuning on R1-Lite and R1-Pro averages 76.7%, versus 53.3% for π0.5 and 24.4% for GR00T-N1.7. The main limitation is sensitivity to low-contrast drawer interiors and incomplete long-term memory, but the results indicate that keeping the VLM as the action-generating agent improves transfer, language grounding, and RL optimization.
Original abstract
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $π_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $π_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.