NTH

Agent as Policy for Robotic Manipulation

AuthorsMengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang

Affiliationsmjia2@nd.edu · University of Notre Dame University of California San Diego San Diego State University*Core contributors

September 21, 2026 2 min read
Watch on YouTube
The one-line take

A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.

Key results

80%
Main-task success coverage

AGP achieved at least this success rate in seven of eight evaluated task configurations.

100%
Dice-flipping success

All evaluated dice-flipping trials succeeded.

29.3%
Repeated-assembly time reduction

Two-pair assembly task time decreased from the first to the fifth execution with accumulated experience.

47.0%
Reasoning latency reduction

Reasoning-and-programming latency decreased from the first to the fifth two-pair assembly execution.

34.3%
Transferred-experience time reduction

Successful GPT-5.6 Terra trials became faster after receiving GPT-6 Astra experience.

What the paper found

Agent as Policy introduces AGP, a framework in which a general-purpose multimodal coding agent acts as the robot policy rather than selecting among separately trained skills. Given video, image, or language instructions and a documented robot interface, the agent requests camera and depth observations, writes Python programs for perception and geometry, computes gripper poses or joint targets, sends motion commands, and revises its programs after physical feedback, with model parameters fixed throughout execution. Using OpenAI’s GPT-6 Astra through Codex and the Mink inverse-kinematics library, AGP controlled real robots across AutoMate-based four-pair assembly, image-guided block construction, dice flipping, targeted throwing, and bimanual towel folding. It achieved at least 80% success in seven of eight main task configurations, including 100% on dice flipping, while simultaneous towel folding reached only 60%, exposing the difficulty of deformable-object coordination. Persistent files containing measurements, corrections, and executable procedures reduced two-pair assembly task time by 29.3% and reasoning-and-programming latency by 47.0% between the first and fifth executions. Experience transfer also raised a weaker GPT-5.6 Terra agent’s success rate from 20% to 80% and reduced successful-trial completion time by 34.3%. A comparison involving Anthropic’s Claude Code and Claude Opus 5 shows that AGP is not tied to OpenAI, but execution latency, inference cost, and reliability remain substantial deployment barriers.

Original abstract

We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis