NTH

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

AuthorsLiyang Li, Muzhi Zhu, Zhiyue Zhao, Hengyu Zhao, Ke Liu, Linhao Zhong, Hao Chen, Chunhua Shen

June 3, 2026 2 min read
Watch on YouTube
The one-line take

This work asks whether foundation models can actively move to match a target view, and shows that new training methods can dramatically improve their ability to navigate in 3D indoor environments.

Key results

500
Evaluation tasks

TVRBench contains 500 evaluation tasks total, split evenly across the four task categories.

7.8%
Strongest open-source success

The strongest open-source model, Qwen3.5-27B, reaches 7.8% success on the evaluation split.

12.0%
Best closed-source success

The strongest closed-source model, Gemini-3.1-Pro, reaches 12.0% success on the evaluation split.

What the paper found

Where to Look introduces Target Viewpoint Reproduction, or TVR, a closed-loop embodied task where a multimodal foundation model must move in a 3D indoor simulator until its egocentric camera exactly matches a target image. The authors, from Zhejiang University, build TVRBench on AI2-THOR and ProcTHOR-10k, spanning 500 evaluation tasks across single-room and multi-room scenes with an exact pose-matching criterion on a 0.25 m and 45°/30° action grid. Evaluation shows a large human gap: the strongest open-source model, Qwen3.5-27B, reaches 7.8% success, the best closed-source model, Gemini-3.1-Pro, reaches 12.0%, while humans score 93.0% on a matched subset. The failure analysis is precise: models over-rotate, rarely issue Stop correctly, and struggle most with body translation, not just visual matching. A unified post-training study finds that supervised fine-tuning on expert visual-action trajectories is the key improvement, lifting Qwen3.5-9B to 50.8% success; adding trajectory-level Multi-turn GRPO on live simulator rollouts nudges this to 51.4%, mainly improving multi-room navigation, whereas chain-of-thought supervision and Single-turn GRPO actually degrade closed-loop performance. The paper argues that active spatial intelligence requires learning perception-to-action mappings over full trajectories, not isolated per-step action prediction.

Original abstract

Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passive understanding of pre-collected observations. We introduce Target Viewpoint Reproduction (TVR) -- an active task where an agent adjusts its viewpoint in a 3D environment until its observation matches a given target image -- and TVRBench, an indoor-simulation benchmark spanning scene scale and target-view visual richness. TVR is far from solved: on the evaluation split, the strongest open-source and closed-source models reach only 7.8% and 12.0% success. Fine-grained analysis identifies two consistent bottlenecks: off-the-shelf models struggle with multi-turn visual history, and performance drops sharply when viewpoint reproduction requires body translation rather than in-place rotation, exposing a gap in mapping spatial discrepancies to embodied movement. To study reducing this gap, we build a unified TVR post-training framework covering expert-trajectory SFT, rationale-supervised CoT-SFT, offline Single-turn GRPO, and on-policy Multi-turn GRPO from live simulator rollouts. Visual-action SFT supplies the main gain, raising a 9B open-source model to 50.8% success; Multi-turn GRPO provides targeted multi-room refinement and reaches 51.4% overall, while CoT supervision and Single-turn GRPO degrade closed-loop performance. These results establish TVRBench as a testbed for measuring and training foundation models that actively perceive and act in 3D environments. Our code, data, and models are available at https://github.com/aim-uofa/TVRBench.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis