Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
AuthorsGabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
AffiliationsTextQL
Resources
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
Key results
Oracle E-Business Suite tables in the simulated enterprise warehouse.
Total rows in the private evaluation warehouse.
End-to-end tasks across five enterprise business areas.
Share of tasks scoring at least 95 points.
Average task score out of 100.
Observed coverage of nominal 80% prediction intervals.
What the paper found
Argo-Bench evaluates whether data agents can complete enterprise workflows rather than merely generate text-to-SQL. It simulates a New York City food-delivery platform, modeled on businesses such as DoorDash, Grubhub, and Uber Eats, and projects its 2024 operations into an Oracle E-Business Suite warehouse with 235 mutually constraining tables and 7.49B rows. The benchmark contains 210 tasks spanning fraud enforcement, forecasting, financial planning, accounting, marketplace decisions, and dashboard publication. Agents receive only an incomplete warehouse view, use SQL and Python with statistical, machine-learning, and optimization libraries, then file actions such as banning accounts, reallocating courier incentives, issuing back pay, or publishing data sources. A simulator retains hidden fraud labels, future outcomes, and counterfactual economics, allowing objective grading by business consequences rather than answer-key agreement or an LLM judge. Across 14 models, Claude Opus 5.5 led with a 59.5 mean score and solved 34.8% of tasks at the 95-point threshold; GPT-6 Astra led forecasting, while Claude Sonnet 5.5 led compliance. Gemini 3.8 Flash, DeepSeek V4.1 Flash, and other models often performed extensive analysis but targeted the wrong record, quantity, or optimization objective. Forecasting was particularly overconfident: nominal 80% intervals covered realized values only 44.8% of the time. The central finding is that enterprise data-agent reliability depends on reconstructing business semantics across tables and aligning actions with operational objectives, not just producing syntactically correct queries.
Original abstract
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.