NTH

Can Predicted Dynamics Exist in the Physical World?

AuthorsBarak Or

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks a simple but important question: can a model’s predicted action or dynamics actually work in the real physical world, and it proposes a gate to reject implausible proposals before execution.

Key results

32
horizon

short action-window and rollout horizon used in the PushT evaluation

0.982
transition-RMSE AUC

strongest scalar detector for dynamic-violation detection on LeRobot PushT

0.972
dynamics residual AUC

standardized dynamics residual detector on LeRobot PushT

0.592
kinematic-only AUC

kinematic-only monitor on LeRobot PushT

0.957
full gate AUC

integrated physical-admissibility gate on LeRobot PushT

87.7%
invalid proposals prevented

full physical gate replay intervention result

What the paper found

This paper asks whether a predicted rollout, action chunk, or latent plan can be physically executable before execution, and argues that low prediction error is not enough. The contribution is a model-agnostic physical-admissibility gate that decomposes executability into flow consistency, recursive reachability, bounded differential growth, and learned dynamics consistency, then rejects any decoded proposal whose worst normalized residual exceeds a threshold. The gate is evaluated on Hugging Face LeRobot PushT, using compact MLP world-model baselines with a horizon of 32 and controlled falsification families including smooth impulse, actuator lag, time warp, mode change, action-state mismatch, and action saturation. On dynamic-violation detection, the strongest scalar residual is the transition-RMSE detector with AUC 0.982 and AP 0.997, followed by the standardized dynamics residual with AUC 0.972 and AP 0.995; kinematic-only scoring is much weaker at AUC 0.592, while the full admissibility gate reaches AUC 0.957 and AP 0.993 with condition-level attribution. In replay intervention, the full physical gate prevents 87.7% of invalid proposals while causing 8.5% false interventions, and retained nominal progress stays near 0.998. The paper also shows that history-conditioned prediction is more accurate than a state-only Markov ensemble on PushT, with rollout RMSE 0.00221 versus 0.01000, exposing partial observability in the monitored state and motivating runtime verification at the prediction-control interface.

Original abstract

Predictive Physical AI systems output state rollouts, action chunks, and latent plans, yet a low root-mean-square error (RMSE) does not imply that a particular proposal is physically executable. We formulate physical admissibility as a prediction-control interface: before execution, a decoded proposal is treated as candidate dynamics and evaluated using kinematic, dynamic, and direct-to-composed horizon conditions. Passing is not a certificate of task success; rejection identifies violation of the specified physical envelope and gives a component-level reason. On Hugging Face LeRobot PushT, controlled falsification shows that one-step prediction-RMSE and standardized dynamics residuals reach area under the receiver operating characteristic curve (AUC) 0.982 and 0.972, kinematic-only conditions reach AUC 0.592, and the full gate reaches AUC 0.957 with condition-level attribution. In replay-based intervention experiments, residual-based filters and the full physical-admissibility gate prevent 87-$89% of invalid proposals while preserving mean progress near 0.998.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis