NTH

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

AuthorsMizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince

August 13, 2026 2 min read
Watch on YouTube
The one-line take

DSAgentBench tests whether AI agents can genuinely complete end-to-end data-science projects on real computers, revealing that current systems still struggle substantially.

Key results

275
Benchmark tasks

Human-authored end-to-end data-science tasks across the full workflow lifecycle.

15
Models evaluated

Closed-source, hybrid, and open-source vision-language agents.

56.70%
Best agent success

Claude-4.6-Sonnet with screenshot-plus-accessibility-tree observations.

85.09%
Human performance

Overall task success under the same deterministic evaluation protocol.

47.6%
Hard tasks

Share of benchmark tasks classified as hard.

What the paper found

DSAgentBench introduces a benchmark for testing whether AI agents can complete end-to-end data-science workflows inside real operating systems rather than merely generate executable code. Extending OSWorld with Jupyter Notebook, VS Code, Chrome, Kaggle API, OpenML, and SQLite, it provides 275 human-authored tasks spanning data acquisition, exploratory analysis, feature engineering, modeling, evaluation, visualization, and reporting. Each task runs in a configured Ubuntu environment and is scored by deterministic evaluators that verify files, numerical results, model performance, visualization metadata, and semantic correctness. Across 15 models, including OpenAI’s GPT-4o and GPT-5, Google’s Gemini 2.5 Pro, Anthropic’s Claude models, and open-source GUI agents, Claude-4.6-Sonnet performs best at 56.70% task success with screenshot-plus-accessibility-tree input, while open-source agents remain below 1%. Human performance reaches 85.09%, exposing a substantial autonomy gap. The benchmark is deliberately long-horizon: 47.6% of tasks are hard, and many require cross-application state management, iterative debugging, and tool orchestration. Ablations show that increasing the action budget provides only marginal gains, indicating that failures arise mainly from UI grounding, planning, reasoning, and analytical execution rather than insufficient interaction length.

Original abstract

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments. Yet existing benchmarks lack real-computer interaction and do not evaluate whether agents can execute complete end-to-end data-science workflows in realistic computing environments, failing to capture the multi-stage, multi-tool nature of data-science practice. We introduce DSAgentBench, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments. DSAgentBench contains 275 diverse tasks covering the entire data-science life-cycle, reflecting the complexity and tool coordination required in practice. Each task requires grounding decisions in intermediate outputs and coordinated tool use, and includes a deterministic evaluator that verifies analytical correctness, visual outputs, and model performance rather than code-only execution. Our extensive experiments with 15 closed- and open-source models show that even the strongest agent, Claude-4.6-Sonnet, achieves only 56.70% task success, while all open-source agents remain below 1%, frequently failing at tool orchestration, OS grounding, and multi-step reasoning. These results reveal a substantial capability gap between current agentic systems and real data-science workflows, positioning DSAgentBench as a foundation for developing grounded, verifiable, autonomous data-science agents. We release DSAgentBench at https://github.com/vis-nlp/DSAgentBench.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis