NTH

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

AuthorsJiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun

July 30, 2026 2 min read
Watch on YouTube
The one-line take

DocOps tests whether AI agents can perform complex document workflows without losing state, misunderstanding content, or damaging document structure.

Key results

210
DocOps task count

Total benchmark tasks spanning atomic, composite, workflow, and cross-document operations.

0.671
Best overall pass rate

GPT-5.5 with OpenAI Codex and document skills.

0.725
GPT-5.5 L1 pass rate

Average performance on atomic-edit tasks.

0.237
GPT-5.5 L4 pass rate

Average performance on cross-document workflow tasks.

95.31%
Verifier human agreement

Agreement across 128 manually audited verifier decisions.

96.67%
Verifier mutation detection

Detected 174 of 180 controlled document violations.

What the paper found

DocOps, from researchers at the Chinese Academy of Sciences and Baidu, is a verifiable benchmark for autonomous agents that manipulate native XLSX, DOCX, PPTX, and PDF files rather than merely reading them. Its taxonomy spans three operation families—content, format, and structure—and four difficulty levels: L1 atomic edits, L2 composite edits, L3 single-document workflows, and L4 cross-document workflows. The benchmark contains 210 tasks, each evaluated by an offline deterministic verifier using structural predicates, linguistic anchors, and preservation checks against the final artifact. Across models and harnesses including OpenAI’s GPT-5.5 with Codex, Anthropic’s Claude Sonnet 4.6 with Claude Code, and open-source systems such as DeepSeek-V4-Pro and Qwen3.6-27B, the best configuration reaches only a 0.671 pass rate. GPT-5.5’s average pass rate falls from 0.725 on L1 tasks to 0.237 on L4 tasks, showing that long-horizon state tracking and cross-document consistency remain major bottlenecks. Excel workflows are especially fragile because agents must preserve formulas, references, validation boundaries, and sheet state. The study identifies three dominant failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing that flattens native structures or breaks formulas and tables. The verifier agrees with human judgments in 95.31% of 128 audited decisions and detects 96.67% of 180 injected violations. Execution harnesses strongly affect results: programmable environments with file-system feedback outperform constrained document APIs, while skills help selectively, improving Qwen3.5-27B under Claude Code by 7.1 percentage points but producing inconsistent gains overall.

Original abstract

As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis