VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
AuthorsHaodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
Resources
VideoCoCo uses executable Blender code as chain-of-thought to plan physical dynamics before a generative video model turns them into realistic videos.
Key results
Draft–instruction–target triplets used to adapt the video editor.
OmniWeaving average physical-consistency score before VideoCoCo.
Average score after adding VideoCoCo to OmniWeaving.
Physical-plausibility average, compared with 52.18% for OmniWeaving.
Point gain in average physical plausibility over OmniWeaving.
What the paper found
VideoCoCo introduces an agentic dual-engine system that treats executable Blender Python as a process-level chain of thought for physically consistent video generation. A coding agent converts a text prompt into a runnable scene-and-dynamics program, Blender executes it in a sandbox to produce a deterministic, temporally dense white-clay draft, and a draft-conditioned video editor transforms that scaffold into photorealistic footage. The authors train the editor with VideoCoCo-3K, a dataset of 3000 draft–instruction–target triplets generated using ByteDance’s Seedance 2.0 as a teacher. Starting from the OmniWeaving generator, VideoCoCo raises the PhyGenBench average from 0.475 to 0.558, outperforming the strongest listed baseline, Wan2.2-TI2V-5B at 0.544, with especially large gains in material and thermal dynamics. On VBench-2.0, the average physical-plausibility score increases from 52.18% to 77.88%, a 25.70-point improvement; mechanics reaches 92.31% and thermotics 72.92%. The ablation shows that executable drafting alone helps, while editor adaptation adds further gains: LoRA tuning reaches 0.558 versus 0.535 for full fine-tuning, suggesting that low-rank adaptation preserves visual priors while learning the draft-to-realistic mapping. Evaluation uses GPT-4o from OpenAI as the PhyGenBench judge and compares against systems including Sora, HunyuanVideo, CogVideoX, and NVIDIA’s Cosmos-Predict2.5. The tradeoff is added inference latency and limited expressiveness for complex phenomena such as turbulent fluids.
Original abstract
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.