Same Weights, Different Robot: A Deployment Safety View of VLA Policies
AuthorsJianwei Tai
Resources
This paper shows that in robot vision-language-action systems, the same checkpoint can behave like a different policy once action normalization metadata changes, creating a hidden safety risk before rollout.
Key results
Mean L2 displacement over six non-gripper action dimensions for LONG - V 2 substitution
Correct metadata at α=0 before substitution
Full substitution with LONG - V 2
Full substitution with LONG - V 2
Full substitution across the four plausible sibling keys
What the paper found
Same Weights, Different Robot argues that for vision-language-action policies, checkpoint equality is not deployment equality: the executable robot policy also includes the action unnormalizer and controller conventions that convert normalized outputs into physical commands. The paper formalizes this as an executable-policy specification, then derives a closed-form affine displacement for quantile-style action normalization metadata, allowing a pre-rollout ExecSpec certificate to measure semantic drift without running inference or rollout. On LIBERO-Goal, substituting the plausible sibling key LONG - V 2 yields mean drift 0.199 over six non-gripper action dimensions, with p95 drift 0.275, and drops replay success from 28/28 to 2/28; on LIBERO-Spatial, the same substitution drops success from 26/26 to 0/26. Across the other suite families, full substitution also gives 0/28 success on all four LIBERO-Object substitutions and 0/23 or 1/23 on LIBERO-Long. The dose-response interpolation from correct to wrong metadata shows success collapsing as the executable specification moves, e.g. Goal with LONG - V 2 goes from 28/28 at α=0 to 17/28, 7/28, 0/28, and 2/28 at α=1. The paper’s deployment-safety claim is narrow but sharp: action-space metadata is part of the policy identity and should be versioned, hashed, and checked before rollout, alongside systems like OpenVLA and LeRobot-style interfaces.
Original abstract
Vision-language-action (VLA) policies are often treated as checkpoint-defined objects: if the weights, prompt, and benchmark suite match, the deployment is assumed to be the same policy. Robot execution breaks this assumption because the same normalized model output can become a different physical action after action unnormalization and controller conventions are applied. This creates a deployment-safety gap: safety review can certify the checkpoint while missing the executable robot policy that reaches the controller. We formalize this gap as an executable policy specification problem: a VLA policy includes the learned model, action representation, metadata-selected unnormalizer, and controller-facing conventions. Under this view, identical checkpoints can be executable-inequivalent. For quantile-style action normalization, we derive a closed-form metadata mismatch transform and an ExecSpec certificate that measures action-space semantic drift without model inference or rollout. On LIBERO-Goal replay, substituting a plausible sibling metadata key yields mean drift 0.199 over six non-gripper action dimensions and reduces success from 28/28 to 2/28 under full substitution. On LIBERO-Spatial replay, the same substituted key reduces success from 26/26 to 0/26. The same full-substitution protocol gives 0/28 success for all four Object substitutions and 0/23 or 1/23 success on Long. Identity-key, replay-validity, no-op filtering, raw-vs-correct replay, mask/gripper, synthetic upper-bound, and OpenVLA-style unnormalizer interface checks rule out several simpler explanations. These results do not certify closed-loop or hardware safety. They support a narrower deployment-safety view: action-space metadata is part of the executable policy and should be checked before rollout.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.