NTH

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

AuthorsYunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Zhuoer Xu, Xiaohu Du, Jianan Ma, Zixing Chen, Xingjun Ma, Yunhao Chen, Xinhao Deng

July 10, 2026 3 min read
Watch on YouTube
The one-line take

This paper presents Vera, a scalable system for discovering and testing safety risks in LLM agents, and releases a new benchmark showing that even state-of-the-art agent frameworks can be surprisingly vulnerable.

Key results

77
Attack methods

leaf-level attack-method taxonomy size

30
Environments

leaf-level execution-environment taxonomy size

39078
Candidate safety cases

combinatorial scenarios produced before filtering

1600
VERA-Bench executable cases

retained executable base scenarios

90.6%
Single-channel ASR

average attack success rate across agents

93.9%
Multi-channel ASR

average attack success rate across agents

What the paper found

V ERA, developed by researchers from AntGroup, Zhejiang University, Fudan University, Alibaba Group, Hunan Institute of Advanced Technology, and Deakin University, is an end-to-end safety testing framework for LLM agents that replaces static red-teaming with executable safety cases, adaptive multi-turn attacks, and evidence-grounded verification over environment state rather than model self-report. The system continuously mines roughly 800 papers to build taxonomies with 124 risk categories, 77 attack methods, and 30 execution environments, then composes them into 39,078 candidate scenarios and releases VERA-Bench with 1,600 executable safety cases across three threat settings. On four production agent frameworks—OpenClaw, Hermes, Codex, and Anthropic’s Claude Code—V ERA exposes severe weaknesses, with average attack success rates of 90.6% in single-channel testing and 93.9% in multi-channel testing, while overall framework results range from 70.3% for OpenClaw to 88.6% for Claude Code. The benchmark also shows that attackability depends sharply on the interaction between risk, environment, and attack method: for example, Integrity averages 95.3% ESR, Harm Output 79.0%, and Social Engineer falls to 42.9% in CRM & Svc. VERA-derived supervision transfers beyond the benchmark, where fine-tuning Qwen3Guard raises R-Judge performance from 59.4% accuracy, 32.3% recall, and 45.8% F1 to 61.7%, 77.9%, and 68.4%, with a separate run reaching 0.930 accuracy and 0.941 F1 and a best evaluation loss of 0.0387 at step 210.

Original abstract

LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis