Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
AuthorsTianxiang Jiang, Linquan Wu, Sheng Xia, Songze Li, Ziang Yan, Haoyu Yang, Yu Qiao, Yi Wang
Resources
The paper teaches video AI to think partly in images instead of only in words, helping it predict future events more accurately.
Key results
Qwen3-VL-8B baseline accuracy on FutureBench
FUTURE-L1-SFT accuracy on FutureBench
FUTURE-L1-RL accuracy on FutureBench
FUTURE-L1-RL exceeds Video-CoE on FutureBench
Qwen3-VL-8B baseline average score on TwiFF-Bench
FUTURE-L1-RL average score on TwiFF-Bench
What the paper found
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction, developed by Shanghai AI Laboratory collaborators with universities including USTC, proposes FUTURE-L1, a Qwen3-VL-8B-based multimodal framework that alternates between text tokens and continuous latent visual spans instead of translating every future hypothesis into language. The core idea is to preserve dynamic future semantics in latent space for video event prediction, then train that behavior in two stages: supervised fine-tuning on FUTURE-L1-50K, a 50K subset curated from TwiFF-2.7M by selecting samples where intermediate future frames provide measurable visual gain, and LA-DAPO reinforcement learning, which adds outcome-contrastive and temporal-diversity rewards to optimize sampled latent trajectories. On FutureBench, FUTURE-L1 lifts Qwen3-VL-8B from 61.0 to 73.2 after SFT and to 85.4 after RL, surpassing Video-CoE by 10.4 points; on TwiFF-Bench, it raises the average score from 2.44 to 3.04. The ablations show the latent MSE weight of 0.1 and latent budget 4 are optimal in SFT, and the best RL coefficients are 0.2 for outcome-contrastive reward and 0.1 for temporal diversity. The paper also reports that FUTURE-L1-RL uses 195.3 tokens and 0.91 seconds per sample on FutureBench, indicating that its accuracy gains come with more compact inference than text-heavy multi-turn video reasoning.
Original abstract
Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in text space: once visual evidence is verbalized, fine-grained motion, geometry, and interaction cues can be lost, leading to plausible but visually ungrounded hallucinations. We introduce Future-L1, an interleaved latent visual reasoning framework that lets an MLLM alternate between language tokens and continuous latent visual spans during autoregressive decoding. To train this capability, we construct Future-L1-50K by selecting examples where future visual hints help prediction and align latent states to future-frame embeddings, then further optimize sampled latent trajectories with LA-DAPO, a latent-aware RL objective with outcome-contrastive and temporal-diversity rewards. Future-L1 achieves new state-of-the-art results on both benchmarks: on FutureBench, it improves Qwen3-VL-8B from 61.0 to 85.4 and exceeds the previous best Video-CoE by 10.4 points; on TwiFF-Bench, it improves the average score from 2.44 to 3.04. These results suggest that future-oriented video reasoning benefits from preserving intermediate visual semantics in latent space rather than translating every reasoning step into text.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.