NTH
AI research

UAM: A Dual-Stream Perspective on Forgetting in VLA Training

AuthorsJianke Zhang, Yuanfei Luo, Yucheng Hu, Xiaoyu Chen, Yanjiang Guo, Ziyang Liu, Hongbin Xu, Tian Lan, Jianyu Chen

May 19, 2026 3 min read
Watch on YouTube
The one-line take

This paper argues that vision-language-action models forget because one encoder is doing two jobs, and fixes it by adding a second brain-like pathway that preserves language understanding while improving robot control.

Key results

3,000 ALOHA trajectories
training trajectories

UAM is trained end-to-end on 3,000 demonstration trajectories collected from the ALOHA bimanual robotic system.

30,000 steps
training steps

The model is trained directly for 30,000 steps on action data only, with no auxiliary vision-language co-training.

over 95%
multimodal retention

After action fine-tuning, UAM retains over 95% of the underlying VLM’s multimodal capability.

less than 5%
average forgetting

The paper reports an embodiment tax of less than 5% on average across multimodal benchmarks.

What the paper found

UAM, or Unified Action Model, argues that standard vision-language-action training imposes an “embodiment tax”: when pretrained vision-language models such as Qwen2.5-7B or PaliGemma are fully fine-tuned on action data, multimodal understanding collapses, with sequential MLP action heads driving VQA scores to 0 and even MoT-coupled models retaining only a fraction of their original capability. The paper measures forgetting as the relative drop in benchmark score across MMMU, MME, MMBench, MathVista, MMStar, and TextVQA, then traces the failure to a representational bottleneck: one encoder is forced to serve both semantic grounding and visuomotor control. UAM breaks that bottleneck by adding a parallel Dorsal Expert, initialized from the generative Bagel checkpoint and trained with a visual-dynamics objective, so the action policy learns through a dual-stream architecture rather than a single shared pathway. Trained end-to-end on 3,000 ALOHA trajectories for 30,000 steps, with no frozen weights, no gradient stopping, and no auxiliary vision-language replay, UAM retains over 95% of the underlying VLM’s multimodal score, with less than a 5% average drop, while also achieving the best average success rate on CALVIN and real-world out-of-distribution manipulation tasks involving unseen objects, novel object-target compositions, and instruction variation. Attention analysis shows a clean functional split: semantic tokens attend to target objects and goal regions, while dorsal tokens focus on the robot arm, interaction boundaries, and scene dynamics, indicating that architectural separation alone can preserve semantic competence while improving control generalization.

Original abstract

Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show that this standard recipe systematically erodes the VLM's multimodal competence, a side effect we call the embodiment tax. But do VLAs have to forget? Inspired by the two-stream organization of biological vision, we trace this degradation to a structural bottleneck: current VLAs ask a single encoder to support both language-grounded semantics and control-relevant visual features, whereas biological vision separates recognition and visuomotor control into distinct pathways. Building on this view, we propose the Unified Action Model (UAM), which adds a parallel Dorsal Expert, an analog of the brain's dorsal pathway. To make the Dorsal Expert an effective second pathway and reduce the control-learning burden on the VLM, we initialize it from a pretrained generative model and train it with a mid-level reasoning objective that predicts visual dynamics. This design allows us to train the whole VLA end-to-end on action data alone: with no parameter freezing, no gradient stopping, and no auxiliary VL co-training, UAM retains over $95\%$ of the underlying VLM's multimodal capability and at the same time achieves the highest average success rate among baselines on a variety of manipulation tasks that probe out-of-distribution generalization, including unseen objects, novel object--target compositions, and instruction variation. Together, these results suggest that semantic preservation in VLAs can emerge from architectural separation itself, rather than being enforced by frozen weights or auxiliary data replay, and that this preserved semantic capability can naturally transfer from VLMs to semantic generalization in actions.

Read the original paper