NTH

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

AuthorsSiyi Chen, Hugo Hadfield, Alex Zook, Mikaela Angelina Uy, Chan Hee Song, Erwin Coumans, Xuning Yang, Faisal Ladhak, Qing Qu, Stan Birchfield, Jonathan Tremblay, Valts Blukis

June 15, 2026 2 min read
Watch on YouTube
The one-line take

VoLo teaches a vision-language model to act like a robot conductor, dynamically steering tools and recovery steps during long-horizon manipulation instead of waiting for each action to finish.

Key results

41.80%
RoboVoLo overall success

VoLoAgent full system on RoboVoLo

12.57%
Pure VLA overall success

Single-action-model baseline on RoboVoLo

17.76%
No-VLA tool-chain success

VoLoAgent ablation without VLA

42.9%
Real-robot success

VoLoAgent full system on 14 real tasks

14.3%
π0.5 real-robot success

Real-robot baseline on the same 14 tasks

What the paper found

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation, from NVIDIA and the University of Michigan, reframes robot control as “physical orchestration”: instead of treating a vision-language-action model as a fixed executor, VoLoAgent uses a VLM to plan, monitor, interrupt, and recover by steering a VLA/WAM, perception models, and grasp/place primitives as callable tools in one closed loop. To benchmark this setting, the paper introduces RoboVoLo, a high-fidelity benchmark built on RoboLab and NVIDIA Isaac Lab with 4 suites, 15 task categories, and 126 tasks spanning common sense, memory, complex references, and world knowledge. On RoboVoLo, the full system reaches 41.80% overall success, versus 12.57% for a pure VLA, 17.76% for a no-VLA tool-chain, and 34.97% for a VLA-only ablation, with the largest gains coming from monitoring and recovery rather than open-loop action. The paper validates the design on a real Franka FR3 cell: full VoLoAgent achieves 42.9% success on 14 real tasks, compared with 14.3% for π0.5. Failure analysis shows that completion-monitor errors dominate across frontier VLMs, while the orchestrator’s recovery loop substantially improves resilience, raising recovery from 13% to 54% on episodes that do fail. The strongest practical takeaway is that long-horizon manipulation improves when the robot can switch between learned policies and perception-grounded primitives mid-execution, rather than committing to a single monolithic action model.

Original abstract

Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning, executing, monitoring, and recovering from failures. We address these demands with a closed agent loop in which a VLM orchestrates heterogeneous robot capabilities as interruptible tools. Unlike in virtual AI agents, the timing of decisions, actions and tool calls is important in a physical world that does not pause for reasoning. We refer to this setting as Physical Orchestration, and propose VoLoAgent, a VLM that plans, monitors, and recovers by treating a VLA/WAM as an interruptible tool it steers mid-rollout alongside vision models and action primitives. To evaluate these long-horizon capabilities, we introduce RoboVoLo, a high-fidelity benchmark for open-vocabulary long-horizon manipulation across common sense, memory/state tracking, complex references, and world knowledge, with both task-level success and failure-mode diagnostics. Experiments show VoLoAgent substantially outperforms single VLA/VLM or tool-based systems, with validation on real-robot experiments. Project page: https://chicychen.github.io/VoLo/

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis