NTH

G0.5: One Autoregressive Stream for Robot Reasoning and Action

AuthorsYicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao

August 19, 2026 2 min read
Watch on YouTube
The one-line take

G0.5 turns a vision-language model into a single-stream robot brain that explains what it is doing while directly generating the actions to do it.

Key results

82.5%
DROID held-out transfer

Average success after DROID post-training on unseen environments and object instances

98.9%
LIBERO success

Overall benchmark success rate

93.3%
RoboTwin 2.0 success

Average success across clean and randomized evaluation

31.4%
BEHAVIOR Challenge score

Task Success Score across 50 household tasks after four post-training epochs

76.7%
Real-world fine-tuning

Average success on R1-Lite and R1-Pro manipulation settings

What the paper found

G0.5 replaces the prevailing VLA design—where a VLM such as Qwen3.5 conditions a separate flow-matching expert—with one autoregressive transformer that reasons and acts in the same token stream. Initialized from Qwen3.5 2B, it uses a learned cross-embodiment ActionCodec with residual vector quantization, a 27-dimensional structured action space, active-degree-of-freedom tokenization, native chain-of-thought for subtasks, bounding boxes, 2D traces, and action hints, plus factorized visual memory for multi-second history. A single next-token cross-entropy objective trains language, reasoning, and motor commands, while exact action-token likelihoods make reinforcement learning with GRPO more direct than for flow matching. Pretraining combines 14 robot embodiments, robot demonstrations, VQA data, and automated annotations generated with Gemini 3, Doubao Seed 2.0 Pro, and SAM3. After post-training, G0.5 reaches 82.5% on DROID with held-out environments and object instances, 98.9% on LIBERO, 93.3% on RoboTwin 2.0, and 87.3% on SimplerEnv-Bridge. On the 2025 BEHAVIOR Challenge, using one generalist checkpoint across 50 household tasks, it achieves 31.4%, exceeding π0.5’s 26.3%; real-world fine-tuning on R1-Lite and R1-Pro averages 76.7%, versus 53.3% for π0.5 and 24.4% for GR00T-N1.7. The main limitation is sensitivity to low-contrast drawer interiors and incomplete long-term memory, but the results indicate that keeping the VLM as the action-generating agent improves transfer, language grounding, and RL optimization.

Original abstract

The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $π_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $π_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis