UniVR: Thinking in Visual Space for Unified Visual Reasoning
AuthorsZhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
Resources
UniVR teaches AI to reason, understand physics, and plan long tasks directly from visual experience without relying on language supervision.
Key results
Raw visual data curated from 16 sources for the benchmark and training pipeline.
Curated supervised examples used to initialize UniVR.
Parameter scale of the Emu3.5-based UniVR model.
UniVR’s overall benchmark score, compared with 39.8 for Emu3.5.
Overall VR-X improvement over the Emu3.5 baseline.
What the paper found
Researchers at Beijing Jiaotong University and ByteDance introduce UniVR, a unified autoregressive model that learns reasoning and planning directly in visual space rather than converting visual states into language. Built from Emu3.5-34B, UniVR uses VR-GRPO, a reinforcement-learning method combining a global task-completion reward with a Step-Focal reward that samples high-uncertainty substeps using CLIP feature variance, targeting logical gaps and physical violations in long trajectories. The team creates VR-X from 1.5M raw samples across 16 sources, retaining 310k cold-start examples, 3k reinforcement-learning examples, and 1.8k evaluation trajectories covering manipulation, cooking, navigation, puzzles, search, editing, and spatial reasoning. On VR-X, UniVR reaches an overall score of 58.2, versus 39.8 for Emu3.5, an 18.4% gain, with improvements reaching 25% across individual categories. It also improves multimodal understanding, raising MMMU from 0.292 to 0.337 and MM-Vet from 28.0 to 35.6. At 34B parameters, UniVR surpasses or approaches pipelines such as Google DeepMind’s Gemini 3 Pro with Nano Banana 2 and competes favorably with OpenAI’s GPT-5, demonstrating that raw visual demonstrations can teach long-horizon policies without dense image-text supervision.
Original abstract
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.