DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
AuthorsJiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
Resources
DocOps tests whether AI agents can perform complex document workflows without losing state, misunderstanding content, or damaging document structure.
Key results
Total benchmark tasks spanning atomic, composite, workflow, and cross-document operations.
GPT-5.5 with OpenAI Codex and document skills.
Average performance on atomic-edit tasks.
Average performance on cross-document workflow tasks.
Agreement across 128 manually audited verifier decisions.
Detected 174 of 180 controlled document violations.
What the paper found
DocOps, from researchers at the Chinese Academy of Sciences and Baidu, is a verifiable benchmark for autonomous agents that manipulate native XLSX, DOCX, PPTX, and PDF files rather than merely reading them. Its taxonomy spans three operation families—content, format, and structure—and four difficulty levels: L1 atomic edits, L2 composite edits, L3 single-document workflows, and L4 cross-document workflows. The benchmark contains 210 tasks, each evaluated by an offline deterministic verifier using structural predicates, linguistic anchors, and preservation checks against the final artifact. Across models and harnesses including OpenAI’s GPT-5.5 with Codex, Anthropic’s Claude Sonnet 4.6 with Claude Code, and open-source systems such as DeepSeek-V4-Pro and Qwen3.6-27B, the best configuration reaches only a 0.671 pass rate. GPT-5.5’s average pass rate falls from 0.725 on L1 tasks to 0.237 on L4 tasks, showing that long-horizon state tracking and cross-document consistency remain major bottlenecks. Excel workflows are especially fragile because agents must preserve formulas, references, validation boundaries, and sheet state. The study identifies three dominant failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing that flattens native structures or breaks formulas and tables. The verifier agrees with human judgments in 95.31% of 128 audited decisions and detects 96.67% of 180 injected violations. Execution harnesses strongly affect results: programmable environments with file-system feedback outperform constrained document APIs, while skills help selectively, improving Qwen3.5-27B under Claude Code by 7.1 percentage points but producing inconsistent gains overall.
Original abstract
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.