NTH

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

AuthorsLehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li

July 30, 2026 2 min read
Watch on YouTube
The one-line take

SecRespond tests whether AI agents can investigate and remediate already-compromised computers, revealing that current systems find known alerts but miss hidden intrusions and produce incomplete recovery plans.

Key results

10
SecRespond benchmark scale

Cyber ranges used to evaluate post-compromise incident response.

23
Evaluation model count

Frontier LLMs evaluated through the OpenCode agent harness.

280
SecRespond rubric

Fine-grained checkpoints mapped to 52 capability items.

79.0%
Claude Opus 4.7 detection

Highest overall range-level detection score.

34.7%
GPT-5.5 detection-planning gap

Difference between GPT-5.5 detection and planning performance.

83%
GPT-5.4 SSH-Miner improvement

Planning score after procedural skills, up from 26%.

What the paper found

SecRespond, developed by Alibaba’s Tongyi Lab, Alibaba Cloud, and The Hong Kong University of Science and Technology, is the first benchmark for post-compromise incident response grounded in forensic disk snapshots rather than isolated alerts or synthetic logs. Agents using the OpenCode harness must reconstruct attacks, verify vulnerabilities and baseline risks, and produce remediation plans across 10 cyber ranges spanning 4 entry-point types, 21 MITRE ATT&CK techniques, and 5 operating systems. Its hierarchical LLM-as-a-Judge framework maps 280 fine-grained checkpoints to 52 capability items covering intrusion entities, persistence, baseline risk, vulnerability risk, and investigation quality. Evaluating 23 models from Anthropic’s Claude, OpenAI’s GPT, Google’s Gemini, DeepSeek, Alibaba’s Qwen, and other families shows a sharp detection-versus-remediation gap: Claude Opus 4.7 leads with 79.0% detection and 65.7% planning, while GPT-5.5 reaches 70.7% detection but only 36.0% planning, a 34.7% gap. Agents usually follow alerted traces but miss silent persistence on disk and stop after obvious fixes, so no model achieves complete detection and remediation on any range. Procedural skills distilled from operational incident-response practice can help: GPT-5.4 improved from 26% to 83% on the SSH-Miner planning task, although skills sometimes narrow investigation and miss long-tail artifacts. The central finding is that current agents can identify visible compromise, but reliable autonomous response still requires systematic host-wide forensics, cross-service attack reconstruction, remediation verification, and business-impact assessment.

Original abstract

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first benchmark for evaluating LLM agents on the post-compromise incident-response workflow. Given a forensic disk snapshot of a compromised host together with the alerts, vulnerability scans, and baseline checks reported by a host security product, agents are required to produce forensic reports on intrusions, baseline risks, and vulnerability risks, together with a remediation plan. We instantiate this task across 10 cyber ranges, each constructed from a distinct compromised cloud host, spanning 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. We evaluate 23 frontier LLMs on the OpenCode agent harness. Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range. This reveals a fundamental bottleneck in building agents for real-world incident response. The benchmark is publicly available at https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis