A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
AuthorsXingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
Resources
The study shows that video models may retain the right physical understanding even after generating the wrong motion—and that targeted internal writes can bring that knowledge back.
Key results
Parameter count of the latent flow-matching video Transformer.
Number of DiT blocks used in the controlled mechanics experiments.
Held-out writes within the near-full recovery window using the boundary-state controller.
Conflict physics-following rate after 10K biased-training steps with neutral-first adaptation.
Physics-following rate from gain-8 amplification of one observed-frame value head across 48 failures.
Parameter scale of the pretrained Wan video DiT used to reproduce writable edits and closure.
What the paper found
This paper asks whether a video model that generates physically incorrect motion has forgotten the correct dynamics or merely failed to use them. In a controlled Spring task, red masses are paired with slow oscillations and blue masses with fast oscillations, creating a shortcut where color can override observed motion. The 488M-parameter model, built from 30 DiT blocks, often follows color, but activation interventions reveal that the motion-consistent future remains causally writable. A compact edit using only 4 principal coordinates, predicted from the boundary position, velocity, and target direction, achieves 85.9% held-out recovery-window success. Writability closes sharply with depth: the same edit changes decoded motion before a commitment boundary but not after it, even though the alternative-motion signal persists. Training later fixes errors that were writable at 3.79 more network sites on average, while longer visible histories and motion-first training delay closure; in a pretrained Wan 1.3B video DiT, neutral-first adaptation reaches 80.5% conflict physics-following at 10K biased-training steps versus 9.4% for direct adaptation. Mechanistically, self-attention transfers motion through observed-frame keys and values into future-frame tokens. Amplifying one value head at gain 8 restores the target motion in 56.3% of 48 failures, whereas excessive gain overshoots. The result separates representation from control: a model may retain a physical solution internally after its ordinary rollout has committed to a shortcut. Replication spans Pendulum, Free Fall, and Wan 1.3B; OpenAI ChatGPT and Codex assisted development, while measurements came from deterministic decoded-video evaluators.
Original abstract
When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.