PANDO: Efficient Multimodal AI Agents via Online Skill Distillation
AuthorsYubo Li, Yidi Miao, Yuntian Shen, Yuxin Liu
Resources
PANDO makes multimodal web agents cheaper over time by distilling useful skills online, cutting token use while improving task success on VisualWebArena.
Key results
Full VisualWebArena benchmark size used for main evaluation
PANDO performance on full VisualWebArena
Strong reproduced baseline on VisualWebArena
Reproduced WALT baseline on VisualWebArena
PANDO mean prompt+completion+reasoning tokens per task
PANDO prompt cache utilization on the full run
What the paper found
PANDO, from Carnegie Mellon University, reframes multimodal web agents around token economics rather than raw rollout scaling: instead of paying more inference on every task, it learns a persistent structured skill library online during VisualWebArena evaluation. The system uses a Plan→Act→Reflect→Learn loop with deterministic keyword-based retrieval, parameterized routines, repeat-loop guardrail rules, confidence-based demotion, polarity-pair merging, hierarchical routing, visual compression, and cache-aware prompting. On the full 910-task VisualWebArena benchmark, PANDO reaches 58.3% success rate, beating SGV at 54.0% and a WALT reproduction at 45.2%, while using 115K tokens per task, 58% fewer than SGV and 61% fewer than WALT. Its efficiency gains show up in intrinsic metrics too: the best automated Action Repetition Rate is 9.1%, Step Overhead Ratio is 1.8×, and prompt cache utilization is 72.4%. A VWA-300 ablation shows that rules and routines supply most of the accuracy lift, while routing, compression, and cache-aware layout convert that larger library into lower marginal cost. Over the stream, the library grows from 12 seed routines to 47 induced routines, with 32 still active after 15 demotions and 11 polarity-pair merges, and later task blocks become cheaper, falling to 103K tokens and 8.9 steps per task by tasks 601–910.
Original abstract
Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raises a central question: can a web agent become more efficient as it accumulates experience, rather than more expensive? We first analyze trajectories from VisualWebArena and identify three recurring sources of inefficiency: repeat-action loops, hidden discovery costs, and low prompt-cache reuse. We then introduce PANDO, a single-rollout online skill-distillation framework that maintains a structured Skill Library and combines progress reflection, confidence-based skill demotion, hierarchical routing, visual compression, and cache-aware prompting. On the full set of 910 VisualWebArena tasks, PANDO achieves a 58.3% success rate, outperforming SGV (54.0%) and our WALT reproduction (45.2%), while using 58% fewer tokens than SGV and 61% fewer tokens than WALT, without any pre-evaluation discovery budget. A 300-task ablation further shows that rules and routines provide most of the success gains, while routing, compression, and cache-aware prompting convert the larger skill library into lower marginal token cost. Finally, we introduce three trajectory-level efficiency metrics -- Action Repetition Rate, Step Overhead Ratio, and Prompt Cache Utilization -- to make efficiency visible beyond terminal success.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.