NTH

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

AuthorsXiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song

July 15, 2026 3 min read
Watch on YouTube
The one-line take

AgentBeats reframes how we evaluate AI agents by standardizing assessment through agent-to-agent protocols, aiming to make benchmarks more open, reproducible, and fair.

Key results

298
Judge agents

Number of judge agents participating in the five-month AgentBeats competition.

94.8%
GPT-5.4 DevEval solve rate

Solve rate for GPT-5.4 with Codex CLI on DevEval.

69.1%
Claude Opus 4.7 SWE-Bench Pro solve rate

Solve rate for Claude Opus 4.7 with Claude Code on SWE-Bench Pro.

5.3
Native versus swapped harness gap

Average percentage-point advantage of native model-harness pairings in the harness-swapping experiment.

What the paper found

A team including Dawn Song at UC Berkeley proposes Agentified Agent Assessment, or AAA, which turns a benchmark into a judge agent and separates assessment logic from the evaluated system through the production-oriented A2A protocol for task management and MCP for tool access. Instead of bespoke benchmark-agent integrations, AAA reduces the architecture to protocol-level compatibility, supports multi-agent evaluation, and offers five AgentBeats deployment modes covering local, remote, hosted, proxy, and continuous-integration workflows. In a five-month competition, AgentBeats attracted 298 judge agents across 12 categories and 467 subject agents, agentifying benchmarks such as Tau2-Bench, MedAgentBench, OfficeQA, OSWorld, and CyberGym. A coding case study evaluated Anthropic’s Claude Opus 4.7 with Claude Code, OpenAI’s GPT-5.4 with Codex CLI, Google DeepMind’s Gemini 3.1 Pro with OpenCode, and Qwen3.5-397B-A17B with mini-SWE-agent on DevEval, SWE-Bench Pro, and Terminal-Bench 2.0. GPT-5.4 with Codex CLI reached a 94.8% solve rate on DevEval, while Claude Opus 4.7 with Claude Code led SWE-Bench Pro at 69.1%. A harness-swapping experiment found native model-harness pairings outperforming swapped pairings by an average of 5.3 percentage points, demonstrating co-adaptation and showing why model-only scores can misrepresent deployed agent performance. Overall, the paper argues that AgentBeats makes agent evaluation more open, interoperable, reproducible, and aligned with production behavior.

Original abstract

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis