NTH

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

AuthorsGabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

AffiliationsTextQL

October 6, 2026 3 min read
Watch on YouTube
The one-line take

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Key results

235
Warehouse tables

Oracle E-Business Suite tables in the simulated enterprise warehouse.

7.49B
Warehouse rows

Total rows in the private evaluation warehouse.

210
Benchmark tasks

End-to-end tasks across five enterprise business areas.

34.8%
Claude Opus 5.5 solved rate

Share of tasks scoring at least 95 points.

59.5
Claude Opus 5.5 mean score

Average task score out of 100.

44.8%
Forecast interval coverage

Observed coverage of nominal 80% prediction intervals.

What the paper found

Argo-Bench evaluates whether data agents can complete enterprise workflows rather than merely generate text-to-SQL. It simulates a New York City food-delivery platform, modeled on businesses such as DoorDash, Grubhub, and Uber Eats, and projects its 2024 operations into an Oracle E-Business Suite warehouse with 235 mutually constraining tables and 7.49B rows. The benchmark contains 210 tasks spanning fraud enforcement, forecasting, financial planning, accounting, marketplace decisions, and dashboard publication. Agents receive only an incomplete warehouse view, use SQL and Python with statistical, machine-learning, and optimization libraries, then file actions such as banning accounts, reallocating courier incentives, issuing back pay, or publishing data sources. A simulator retains hidden fraud labels, future outcomes, and counterfactual economics, allowing objective grading by business consequences rather than answer-key agreement or an LLM judge. Across 14 models, Claude Opus 5.5 led with a 59.5 mean score and solved 34.8% of tasks at the 95-point threshold; GPT-6 Astra led forecasting, while Claude Sonnet 5.5 led compliance. Gemini 3.8 Flash, DeepSeek V4.1 Flash, and other models often performed extensive analysis but targeted the wrong record, quantity, or optimization objective. Forecasting was particularly overconfident: nominal 80% intervals covered realized values only 44.8% of the time. The central finding is that enterprise data-agent reliability depends on reconstructing business semantics across tables and aligning actions with operational objectives, not just producing syntactically correct queries.

Original abstract

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis
03Benchmark

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan

GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.

Read analysis