NTH

UniVR: Thinking in Visual Space for Unified Visual Reasoning

AuthorsZhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin

July 21, 2026 2 min read
Watch on YouTube
The one-line take

UniVR teaches AI to reason, understand physics, and plan long tasks directly from visual experience without relying on language supervision.

Key results

1.5M
VR-X raw samples

Raw visual data curated from 16 sources for the benchmark and training pipeline.

310k
Cold-start samples

Curated supervised examples used to initialize UniVR.

34B
UniVR model size

Parameter scale of the Emu3.5-based UniVR model.

58.2
VR-X overall score

UniVR’s overall benchmark score, compared with 39.8 for Emu3.5.

18.4%
Gain over Emu3.5

Overall VR-X improvement over the Emu3.5 baseline.

What the paper found

Researchers at Beijing Jiaotong University and ByteDance introduce UniVR, a unified autoregressive model that learns reasoning and planning directly in visual space rather than converting visual states into language. Built from Emu3.5-34B, UniVR uses VR-GRPO, a reinforcement-learning method combining a global task-completion reward with a Step-Focal reward that samples high-uncertainty substeps using CLIP feature variance, targeting logical gaps and physical violations in long trajectories. The team creates VR-X from 1.5M raw samples across 16 sources, retaining 310k cold-start examples, 3k reinforcement-learning examples, and 1.8k evaluation trajectories covering manipulation, cooking, navigation, puzzles, search, editing, and spatial reasoning. On VR-X, UniVR reaches an overall score of 58.2, versus 39.8 for Emu3.5, an 18.4% gain, with improvements reaching 25% across individual categories. It also improves multimodal understanding, raising MMMU from 0.292 to 0.337 and MM-Vet from 28.0 to 35.6. At 34B parameters, UniVR surpasses or approaches pipelines such as Google DeepMind’s Gemini 3 Pro with Nano Banana 2 and competes favorably with OpenAI’s GPT-5, demonstrating that raw visual demonstrations can teach long-horizon policies without dense image-text supervision.

Original abstract

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis