Streaming Video Generation with Streaming Force Control
AuthorsHanhui Wang, Yiming Xie, Haiwen Feng, Zhaoyang Lv, Shenlong Wang, Huaizu Jiang
Resources
StreamForce is a real-time video generation system that can instantly react to continuous force inputs, aiming to make generated motion both physically responsive and visually realistic.
Key results
Synthetic force-conditioned videos used for bidirectional teacher training
Pexels-based image-force data used for causal distillation
Frames per second on a single H200 GPU
Output resolution reported for streaming inference
Best reported total score on local-force evaluation
What the paper found
Streaming Video Generation with Streaming Force Control, by Hanhui Wang, Yiming Xie, Haiwen Feng, Zhaoyang Lv, Shenlong Wang, and Huaizu Jiang, introduces StreamForce, a causal video generation system built to respond to continuous physical force inputs in real time rather than offline trajectory prompts. The core idea is a unified force representation that encodes both global forces like wind and local contact forces like pushes in a single pixel-aligned masked tensor, paired with a force-aware distillation pipeline that transfers force-motion behavior from a bidirectional Wan2.2 5B TI2V teacher into an autoregressive student. Training uses 30K synthetic force-conditioned videos, another 30K Blender-rendered force-changing clips, and about 90K diverse Pexels image-force pairs. On a single H200 GPU, the final model streams at 16.6 FPS at 832×480 with 0.6-second latency. In user studies on 40 test cases, StreamForce outperforms baselines on force-changing global and local control, reaching 86.5% force adherence for global changes and 80.4% for local changes, while its Physics-IQ score rises to 46.31 on local-force evaluation and 40.99 on global-force evaluation. The paper’s novelty is not just controllability, but causal responsiveness: users can change force magnitude or direction mid-generation and the model adapts without breaking photometric stability, while also showing emergent intuitive physics such as mass- and friction-dependent motion.
Original abstract
We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train separate models for different force types, assume fixed forces, or rely on non-causal processing, StreamForce is a causal and unified model that responds instantly and coherently to both local and global, time-varying forces. To achieve this, we design a unified force representation as a control signal and develop a distillation pipeline for force-controllable video generation. Our model combines autoregressive efficiency with force responsiveness, sustaining stable photometric and dynamic realism. StreamForce runs at up to 16.6 FPS on a single GPU, achieving state-of-the-art performance in both force adherence and motion realism. Project website: https://neu-vi.github.io/StreamForce/
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.