Agent as Policy for Robotic Manipulation
AuthorsMengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang
Affiliationsmjia2@nd.edu · University of Notre Dame University of California San Diego San Diego State University*Core contributors
Resources
A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.
Key results
AGP achieved at least this success rate in seven of eight evaluated task configurations.
All evaluated dice-flipping trials succeeded.
Two-pair assembly task time decreased from the first to the fifth execution with accumulated experience.
Reasoning-and-programming latency decreased from the first to the fifth two-pair assembly execution.
Successful GPT-5.6 Terra trials became faster after receiving GPT-6 Astra experience.
What the paper found
Agent as Policy introduces AGP, a framework in which a general-purpose multimodal coding agent acts as the robot policy rather than selecting among separately trained skills. Given video, image, or language instructions and a documented robot interface, the agent requests camera and depth observations, writes Python programs for perception and geometry, computes gripper poses or joint targets, sends motion commands, and revises its programs after physical feedback, with model parameters fixed throughout execution. Using OpenAI’s GPT-6 Astra through Codex and the Mink inverse-kinematics library, AGP controlled real robots across AutoMate-based four-pair assembly, image-guided block construction, dice flipping, targeted throwing, and bimanual towel folding. It achieved at least 80% success in seven of eight main task configurations, including 100% on dice flipping, while simultaneous towel folding reached only 60%, exposing the difficulty of deformable-object coordination. Persistent files containing measurements, corrections, and executable procedures reduced two-pair assembly task time by 29.3% and reasoning-and-programming latency by 47.0% between the first and fifth executions. Experience transfer also raised a weaker GPT-5.6 Terra agent’s success rate from 20% to 80% and reduced successful-trial completion time by 34.3%. A comparison involving Anthropic’s Claude Code and Claude Opus 5 shows that AGP is not tied to OpenAI, but execution latency, inference cost, and reliability remain substantial deployment barriers.
Original abstract
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.