NTH

Same Weights, Different Robot: A Deployment Safety View of VLA Policies

AuthorsJianwei Tai

June 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that in robot vision-language-action systems, the same checkpoint can behave like a different policy once action normalization metadata changes, creating a hidden safety risk before rollout.

Key results

0.199
LIBERO-Goal mean drift

Mean L2 displacement over six non-gripper action dimensions for LONG - V 2 substitution

28/28
LIBERO-Goal replay success

Correct metadata at α=0 before substitution

2/28
LIBERO-Goal replay success

Full substitution with LONG - V 2

0/26
LIBERO-Spatial replay success

Full substitution with LONG - V 2

0/28
LIBERO-Object replay success

Full substitution across the four plausible sibling keys

What the paper found

Same Weights, Different Robot argues that for vision-language-action policies, checkpoint equality is not deployment equality: the executable robot policy also includes the action unnormalizer and controller conventions that convert normalized outputs into physical commands. The paper formalizes this as an executable-policy specification, then derives a closed-form affine displacement for quantile-style action normalization metadata, allowing a pre-rollout ExecSpec certificate to measure semantic drift without running inference or rollout. On LIBERO-Goal, substituting the plausible sibling key LONG - V 2 yields mean drift 0.199 over six non-gripper action dimensions, with p95 drift 0.275, and drops replay success from 28/28 to 2/28; on LIBERO-Spatial, the same substitution drops success from 26/26 to 0/26. Across the other suite families, full substitution also gives 0/28 success on all four LIBERO-Object substitutions and 0/23 or 1/23 on LIBERO-Long. The dose-response interpolation from correct to wrong metadata shows success collapsing as the executable specification moves, e.g. Goal with LONG - V 2 goes from 28/28 at α=0 to 17/28, 7/28, 0/28, and 2/28 at α=1. The paper’s deployment-safety claim is narrow but sharp: action-space metadata is part of the policy identity and should be versioned, hashed, and checked before rollout, alongside systems like OpenVLA and LeRobot-style interfaces.

Original abstract

Vision-language-action (VLA) policies are often treated as checkpoint-defined objects: if the weights, prompt, and benchmark suite match, the deployment is assumed to be the same policy. Robot execution breaks this assumption because the same normalized model output can become a different physical action after action unnormalization and controller conventions are applied. This creates a deployment-safety gap: safety review can certify the checkpoint while missing the executable robot policy that reaches the controller. We formalize this gap as an executable policy specification problem: a VLA policy includes the learned model, action representation, metadata-selected unnormalizer, and controller-facing conventions. Under this view, identical checkpoints can be executable-inequivalent. For quantile-style action normalization, we derive a closed-form metadata mismatch transform and an ExecSpec certificate that measures action-space semantic drift without model inference or rollout. On LIBERO-Goal replay, substituting a plausible sibling metadata key yields mean drift 0.199 over six non-gripper action dimensions and reduces success from 28/28 to 2/28 under full substitution. On LIBERO-Spatial replay, the same substituted key reduces success from 26/26 to 0/26. The same full-substitution protocol gives 0/28 success for all four Object substitutions and 0/23 or 1/23 success on Long. Identity-key, replay-validity, no-op filtering, raw-vs-correct replay, mask/gripper, synthetic upper-bound, and OpenVLA-style unnormalizer interface checks rule out several simpler explanations. These results do not certify closed-loop or hardware safety. They support a narrower deployment-safety view: action-space metadata is part of the executable policy and should be checked before rollout.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis