On the Identifiability of Controlled World Models
AuthorsXiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
Resources
This paper explains when learned world models can truly recover both the hidden state and the effects of actions, especially when training data lacks sufficient action diversity.
Key results
One-step transitions used in the representation and action-coverage experiments.
Hidden width of the 8-layer SiLU MLP encoder and action-conditioned predictor.
Constant-action sequences evaluated by the goal-conditioned planning probe.
What the paper found
Researchers from Tsinghua University, Didi Voyager Labs, and the University of Hong Kong study when an action-conditioned world model can recover both latent state and controlled dynamics from nonlinear observations. Building on Yann LeCun’s JEPA family, including LeJEPA, V-JEPA 2, and LeWorldModel, they analyze a stationary linear-Gaussian latent system trained with a LeJEPA-style predictive objective and Gaussian representation constraint. The central result is a joint identifiability theorem: a positive predictable-signal spectral margin, gamma-rep, prevents nonlinear representations from competing with the true latent coordinates, while positive conditional action excitation, rho-tr, identifies the response to every action. Under these conditions, every global optimum recovers the latent state and controlled conditional-mean transition up to a common orthogonal transformation. Approximate identification worsens inversely with gamma-rep, and counterfactual transition error can amplify inversely with rho-tr; for exploration noise level sigma, the amplification factor is 1/sigma squared, diverging when sigma equals 0, even if on-policy prediction remains accurate. Experiments use 100,000 one-step transitions, 8-layer SiLU MLPs with hidden width 512, and four nonlinear observation maps across five independent runs. A planning probe evaluates 210 constant-action sequences and shows that increasing conditional action coverage improves counterfactual prediction and reduces terminal planning error. The paper’s main practical message is that action conditioning alone is insufficient: world-model data must vary actions conditional on state.
Original abstract
Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework for learning such models in representation space. Recent action-conditioned extensions perform promisingly in visual control and latent-space planning, but leave a fundamental question unresolved: when does controlled latent prediction identify both the underlying state and the controlled dynamics? This is challenging under nonlinear observations and behavior policies with limited conditional action variation, where state-dependent evolution and action effects can be statistically confounded. We establish a joint identifiability theory for controlled world models with Gaussian latent states under state-dependent Gaussian behavior policies. We identify two policy-dependent conditions: spectral separation of the predictable signal governs representation identifiability, while non-degenerate conditional action variation governs transition identifiability. We prove that when both conditions hold, every global minimizer of the JEPA objective identifies the latent state and controlled transition up to an orthogonal transformation. We further derive quantitative bounds on representation and transition identifiability under approximate optimization. Finally, we construct predictor perturbations along weakly excited action directions whose counterfactual-to-on-policy error ratio is the inverse transition-identifiability margin, revealing the cost of limited action coverage. Experiments across nonlinear observation maps and behavior policies corroborate the theory and demonstrate implications for transition identifiability, counterfactual prediction, and goal-conditioned latent planning.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.