NTH

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

AuthorsQuan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres

September 7, 2026 2 min read
Watch on YouTube
The one-line take

τ^τ-Bench tests whether AI coding agents can build production-ready customer-service agents, revealing that even strong models succeed on fewer than one-quarter of realistic simulations.

Key results

23.9%
Best developer-agent score

Claude Opus 5 under Claude Code’s pass rate across the 53-task benchmark

82.2%
Expert reference ceiling

Pass rate of expert-authored reference agents

53
Benchmark tasks

Release tasks spanning airline, retail, telecom, and banking

2,868
Evidence artifacts

Distinct multimodal artifacts across the four domains

92%
Single-tool-loop architectures

Share of constructed agents using a single LLM tool loop

67%
Telecom architecture hint gain

Score after adding intent routing and tool-call review, up from 31%

What the paper found

ττ-bench evaluates whether coding agents can construct deployable customer-service agents rather than merely operate prebuilt ones. Each task supplies scattered multimodal business evidence, an interactive client holding undocumented requirements, a potentially defective production REST API, inherited code, and fixed model and serving-cost constraints; the constructed agent is then deployed against held-out simulated customers. Across 53 tasks in airline, retail, telecom, and banking, Claude Opus 5 under Anthropic’s Claude Code achieved the strongest result at 23.9%, versus an expert-authored reference ceiling of 82.2%. The benchmark contains 2,868 evidence artifacts and over 5.5M tokens, forcing systems to reconcile documents, transcripts, spreadsheets, screenshots, recordings, and conflicting records. Agents usually defaulted to a single LLM tool loop—92% of builds—while searching instead of reading, rarely interviewing the client, rewriting inherited systems without testing them, mishandling quiet API defects, and under-optimizing model budgets. On client-enabled tasks, builds asking four or more questions averaged 0.50, compared with 0.16 for builds asking none. A controlled telecom intervention shows that a one-line design hint—route by intent and review tool calls—raised performance from 31% to 67%, indicating that architecture selection and requirements elicitation, not only code generation, are major bottlenecks. The roster also included OpenAI’s GPT-5.6-sol and Google’s Gemini models, highlighting that current frontier coding systems can produce runnable agents but remain far from reliable end-to-end engineering.

Original abstract

LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce $τ^τ$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for $τ^τ$-bench to turn the work of cooperative agent building into a measurable target for coding agents.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis