NTH
AI research

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

AuthorsTao Yu, Hao Wang, Changyu Li, Shenghua Chai, Minghui Zhang, Zhongtian Luo, Yuxuan Zhou, Haopeng Jin, Zhaolu Kang, Jiabing Yang, YiFan Zhang, Xinming Wang, Hongzhu Yi, Zheqi He, Jing-Shu Zheng, Xi Yang, Yan Huang, Liang Wang

May 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a realistic benchmark that tests whether multiple AI agents can actually work together inside enterprise-style organizations with roles, permissions, and policy approvals.

Key results

300 tasks
Benchmark dataset size

EntCollabBench released dataset includes 300 total tasks spanning workflow and approval settings.

160 workflow, 40 workflow multi-task, 80 approval, and 20 approval multi-task
Workflow / approval split

The benchmark is divided into workflow and approval subsets, with these four task counts reported in the dataset statistics.

11 role-specialized agents across six departments
Agent / department count

The simulated organization uses 11 specialized agents distributed across six departments to enforce role isolation and delegation.

290 rules
Policy schema size

The approval subset is built from a structured policy schema extracted from the GitLab Handbook and GDPR articles, finalized to 290 rules.

96.0% and 98.0%
Human agreement

The three-model majority-vote judge matched human annotations in 48/50 cases for Gemini-3.1-Pro-Preview and 49/50 for Qwen3.5-122B-A10B.

62.00% average accuracy
Best overall model accuracy

DeepSeek-V4-Pro achieved the best overall task-level average accuracy on EntCollabBench.

What the paper found

The paper introduces EntCollabBench, a benchmark that tests whether LLM agents can collaborate inside an enterprise organization with role isolation, permission-restricted tools, and stateful workflows, instead of assuming one omnipotent agent with full access. The benchmark simulates 11 specialized agents across six departments and splits evaluation into a Workflow track, where agents must change persistent business-system state, and an Approval track, where finance, legal, and procurement agents make deterministic policy decisions. Workflow instances are generated from 20 domain templates over systems such as ITSM, HR, CSM, Gitea, Email, Calendar, Teams, and Drive, while approval instances are derived from a structured policy schema extracted from 60 pages of the GitLab Handbook plus 11 GDPR articles, yielding 290 rules. The released dataset contains 300 tasks, including 160 workflow, 40 workflow multi-task, 80 approval, and 20 approval multi-task cases. Evaluation uses execution traces, database snapshot diffs, and a three-model majority-vote judge with Gemini-3.1-Pro, GPT-5.4, and Claude-Sonnet-4.6; human agreement is 96.0% and 98.0% on sampled checks. Results show that role-level competence does not translate into end-to-end collaboration: the best overall model, DeepSeek-V4-Pro, reaches 62.00% average accuracy, but its workflow multi-task task success is only 50.00% despite 78.33% subtask accuracy. The main failure modes are delegation omissions, context loss across handoffs, parameter-semantic errors such as wrong enum values or object IDs, and weak decision commitment in approvals, showing that enterprise agent robustness depends on routing and grounding, not just tool use.

Original abstract

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.

Read the original paper