JEPA-Anything: Learning Predictive Models across Different Worlds
AuthorsTaoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu, Zhaochen Yu, Yuying Zhang, Qiang Gao, Mengyue Yang, Wanli Ouyang, Pheng Ann Heng, Yingcheng Wu, Zhenfei Yin, Ling Yang
AffiliationsCorresponding authors. Organizations and contact details are listed in Appendix C
JEPA-Anything aims to provide one factorized recipe for learning predictive world models spanning vision, biology, medicine, physics, control, and weather.
Key results
JEPA-Anything improved reported metrics across all 10 tasks.
Reduction in one-step four-channel MSE on CITRIS Interventional Pong.
Fitted slope for the recovered Keplerian frequency–semimajor-axis relation.
What the paper found
JEPA-Anything proposes orthogonal predictive factorization, or OPF, as a shared world-modeling core for seven domains: vision, biology, clinical forecasting, control, molecular dynamics, physical fields, and weather. Instead of predicting one monolithic JEPA embedding, learned projectors divide the target into complementary orthogonal subspaces, dedicated predictors estimate each factor, activity regularization prevents inactive branches, and a pseudoinverse synthesizes a complete latent state for decoding, intervention prediction, planning, or autoregressive rollout. Using backbones including DINOv3, SigLIP2, scGPT, and GPT-2 small, the method improves matched metrics across 10 dynamics tasks, cuts single-intervention error on CITRIS Interventional Pong by 34.83%, and achieves the lowest one-step and 100-step molecular errors across four systems. It also supports scientific analysis: factor coordinates nominated IL-18 plus CD73 blockade, which showed stronger antitumor activity in co-cultures, organoids, tumor fragments, and mice, while orbital modes recovered Keplerian dynamics with a fitted slope of -1.4991. The work positions OPF as a general alternative to domain-specific latent models, complementing systems such as V-JEPA 2 and DeepMind Control Suite, while noting that predictive factors are not automatically causal. The paper also discloses OpenAI Codex for editing and technical checks.
Original abstract
World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic framework based on orthogonal predictive factorization (OPF). Extending joint-embedding predictive architectures, OPF decomposes latent targets into complementary factors, learns them through dedicated pathways, and recombines them within a shared predictive design. We evaluate JEPA-Anything across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather. Experiments span representation learning, intervention prediction, out-of-distribution generalization, and long-horizon dynamics, including 10 matched dynamics tasks, forecasting of over 1,000 clinical events, and 100-step molecular rollouts across four systems. Against matched JEPA baselines, JEPA-Anything improves reported metrics on all 10 dynamics tasks and reduces single-intervention prediction error on Interventional Pong by 34.8%. It achieves the lowest one-step and 100-step molecular errors among compared methods in all four systems. Beyond prediction, a factor-nominated biological intervention receives experimental support in cell co-cultures, patient-derived organoids, tumor fragments, and mice; latent orbital modes recover the Keplerian scaling exponent with a fitted slope of -1.4991. These results support a common factorized predictive principle across heterogeneous worlds, connecting world modeling with intervention and experimentally grounded scientific discovery. Code: https://github.com/Gen-Verse/JEPA-Anything
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.