Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
AuthorsLiyang Li, Muzhi Zhu, Zhiyue Zhao, Hengyu Zhao, Ke Liu, Linhao Zhong, Hao Chen, Chunhua Shen
This work asks whether foundation models can actively move to match a target view, and shows that new training methods can dramatically improve their ability to navigate in 3D indoor environments.
Key results
TVRBench contains 500 evaluation tasks total, split evenly across the four task categories.
The strongest open-source model, Qwen3.5-27B, reaches 7.8% success on the evaluation split.
The strongest closed-source model, Gemini-3.1-Pro, reaches 12.0% success on the evaluation split.
What the paper found
Where to Look introduces Target Viewpoint Reproduction, or TVR, a closed-loop embodied task where a multimodal foundation model must move in a 3D indoor simulator until its egocentric camera exactly matches a target image. The authors, from Zhejiang University, build TVRBench on AI2-THOR and ProcTHOR-10k, spanning 500 evaluation tasks across single-room and multi-room scenes with an exact pose-matching criterion on a 0.25 m and 45°/30° action grid. Evaluation shows a large human gap: the strongest open-source model, Qwen3.5-27B, reaches 7.8% success, the best closed-source model, Gemini-3.1-Pro, reaches 12.0%, while humans score 93.0% on a matched subset. The failure analysis is precise: models over-rotate, rarely issue Stop correctly, and struggle most with body translation, not just visual matching. A unified post-training study finds that supervised fine-tuning on expert visual-action trajectories is the key improvement, lifting Qwen3.5-9B to 50.8% success; adding trajectory-level Multi-turn GRPO on live simulator rollouts nudges this to 51.4%, mainly improving multi-room navigation, whereas chain-of-thought supervision and Single-turn GRPO actually degrade closed-loop performance. The paper argues that active spatial intelligence requires learning perception-to-action mappings over full trajectories, not isolated per-step action prediction.
Original abstract
Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passive understanding of pre-collected observations. We introduce Target Viewpoint Reproduction (TVR) -- an active task where an agent adjusts its viewpoint in a 3D environment until its observation matches a given target image -- and TVRBench, an indoor-simulation benchmark spanning scene scale and target-view visual richness. TVR is far from solved: on the evaluation split, the strongest open-source and closed-source models reach only 7.8% and 12.0% success. Fine-grained analysis identifies two consistent bottlenecks: off-the-shelf models struggle with multi-turn visual history, and performance drops sharply when viewpoint reproduction requires body translation rather than in-place rotation, exposing a gap in mapping spatial discrepancies to embodied movement. To study reducing this gap, we build a unified TVR post-training framework covering expert-trajectory SFT, rationale-supervised CoT-SFT, offline Single-turn GRPO, and on-policy Multi-turn GRPO from live simulator rollouts. Visual-action SFT supplies the main gain, raising a 9B open-source model to 50.8% success; Multi-turn GRPO provides targeted multi-room refinement and reaches 51.4% overall, while CoT supervision and Single-turn GRPO degrade closed-loop performance. These results establish TVRBench as a testbed for measuring and training foundation models that actively perceive and act in 3D environments. Our code, data, and models are available at https://github.com/aim-uofa/TVRBench.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.