NTH

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation

AuthorsYubo Li, Yidi Miao, Yuntian Shen, Yuxin Liu

July 1, 2026 2 min read
Watch on YouTube
The one-line take

PANDO makes multimodal web agents cheaper over time by distilling useful skills online, cutting token use while improving task success on VisualWebArena.

Key results

910
VWA tasks

Full VisualWebArena benchmark size used for main evaluation

58.3%
success rate

PANDO performance on full VisualWebArena

54.0%
SGV success rate

Strong reproduced baseline on VisualWebArena

45.2%
WALT success rate

Reproduced WALT baseline on VisualWebArena

115K
tokens per task

PANDO mean prompt+completion+reasoning tokens per task

72.4%
cache utilization

PANDO prompt cache utilization on the full run

What the paper found

PANDO, from Carnegie Mellon University, reframes multimodal web agents around token economics rather than raw rollout scaling: instead of paying more inference on every task, it learns a persistent structured skill library online during VisualWebArena evaluation. The system uses a Plan→Act→Reflect→Learn loop with deterministic keyword-based retrieval, parameterized routines, repeat-loop guardrail rules, confidence-based demotion, polarity-pair merging, hierarchical routing, visual compression, and cache-aware prompting. On the full 910-task VisualWebArena benchmark, PANDO reaches 58.3% success rate, beating SGV at 54.0% and a WALT reproduction at 45.2%, while using 115K tokens per task, 58% fewer than SGV and 61% fewer than WALT. Its efficiency gains show up in intrinsic metrics too: the best automated Action Repetition Rate is 9.1%, Step Overhead Ratio is 1.8×, and prompt cache utilization is 72.4%. A VWA-300 ablation shows that rules and routines supply most of the accuracy lift, while routing, compression, and cache-aware layout convert that larger library into lower marginal cost. Over the stream, the library grows from 12 seed routines to 47 induced routines, with 32 still active after 15 demotions and 11 polarity-pair merges, and later task blocks become cheaper, falling to 103K tokens and 8.9 steps per task by tasks 601–910.

Original abstract

Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raises a central question: can a web agent become more efficient as it accumulates experience, rather than more expensive? We first analyze trajectories from VisualWebArena and identify three recurring sources of inefficiency: repeat-action loops, hidden discovery costs, and low prompt-cache reuse. We then introduce PANDO, a single-rollout online skill-distillation framework that maintains a structured Skill Library and combines progress reflection, confidence-based skill demotion, hierarchical routing, visual compression, and cache-aware prompting. On the full set of 910 VisualWebArena tasks, PANDO achieves a 58.3% success rate, outperforming SGV (54.0%) and our WALT reproduction (45.2%), while using 58% fewer tokens than SGV and 61% fewer tokens than WALT, without any pre-evaluation discovery budget. A 300-task ablation further shows that rules and routines provide most of the success gains, while routing, compression, and cache-aware prompting convert the larger skill library into lower marginal token cost. Finally, we introduce three trajectory-level efficiency metrics -- Action Repetition Rate, Step Overhead Ratio, and Prompt Cache Utilization -- to make efficiency visible beyond terminal success.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis