NTH

Agents' Last Exam

AuthorsYiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wa...

June 23, 2026 2 min read
Watch on YouTube
The one-line take

This paper launches Agents' Last Exam, a living benchmark aimed at measuring whether AI agents can actually handle real-world professional tasks that matter economically.

Key results

960
task_workflows

Expert-authored workflows in ALE

1490
task_instances

Runnable benchmark instances

55
subdomains

Workflow-level subdomains in the taxonomy

13
industry_clusters

Top-level industry clusters

250+
expert_contributors

Domain experts involved in development

24.0%
best_overall_pass_rate

Best overall public-set pass rate reported for Codex with GPT-5.5

What the paper found

Agents’ Last Exam, led by the University of California, Berkeley with major contributors from Anthropic, OpenAI, Google, DeepSeek, Meta, and others, introduces a benchmark for generalist computer-use agents that measures long-horizon, economically valuable work rather than short-form question answering. ALE is built from 960 expert-authored task workflows and 1,490 runnable instances spanning 55 subdomains across 13 industry clusters, grounded in SOC 2018 and O*NET and sourced from 250+ domain experts. The benchmark’s core novelty is deterministic verification for heterogeneous deliverables, using code-based scoring, structured rubrics, and task-specific VM verifiers instead of open-ended human judgment. It targets Generalist CUA agents such as Claude Code, Codex, and OpenClaw extended with GUI-as-Tool, and evaluates them on a three-tier public split: Near-Term, Full-Spectrum, and Last-Exam. Results show the hardest tier is still far from solved: the strongest configuration, Codex with GPT-5.5, reaches 38.1% pass rate on Near-Term, 22.7% on Full-Spectrum, and 0.0% on Last-Exam, while the best overall public-set pass rate is only 24.0%. The paper also reports that code-based evaluation covers 93.2% of workflows, LLM-as-judge only 6.8%, and that model choice matters more than harness choice, with a 16.8-point spread across backbones versus 4.9 to 7.2 points across harnesses under fixed models.

Original abstract

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 subfields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is 2.6%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis