Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
AuthorsYunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Zhuoer Xu, Xiaohu Du, Jianan Ma, Zixing Chen, Xingjun Ma, Yunhao Chen, Xinhao Deng
This paper presents Vera, a scalable system for discovering and testing safety risks in LLM agents, and releases a new benchmark showing that even state-of-the-art agent frameworks can be surprisingly vulnerable.
Key results
leaf-level attack-method taxonomy size
leaf-level execution-environment taxonomy size
combinatorial scenarios produced before filtering
retained executable base scenarios
average attack success rate across agents
average attack success rate across agents
What the paper found
V ERA, developed by researchers from AntGroup, Zhejiang University, Fudan University, Alibaba Group, Hunan Institute of Advanced Technology, and Deakin University, is an end-to-end safety testing framework for LLM agents that replaces static red-teaming with executable safety cases, adaptive multi-turn attacks, and evidence-grounded verification over environment state rather than model self-report. The system continuously mines roughly 800 papers to build taxonomies with 124 risk categories, 77 attack methods, and 30 execution environments, then composes them into 39,078 candidate scenarios and releases VERA-Bench with 1,600 executable safety cases across three threat settings. On four production agent frameworks—OpenClaw, Hermes, Codex, and Anthropic’s Claude Code—V ERA exposes severe weaknesses, with average attack success rates of 90.6% in single-channel testing and 93.9% in multi-channel testing, while overall framework results range from 70.3% for OpenClaw to 88.6% for Claude Code. The benchmark also shows that attackability depends sharply on the interaction between risk, environment, and attack method: for example, Integrity averages 95.3% ESR, Harm Output 79.0%, and Social Engineer falls to 42.9% in CRM & Svc. VERA-derived supervision transfers beyond the benchmark, where fine-tuning Qwen3Guard raises R-Judge performance from 59.4% accuracy, 32.3% recall, and 45.8% F1 to 61.7%, 77.9%, and 68.4%, with a separate run reaching 0.930 accuracy and 0.941 F1 and a best evaluation loss of 0.0387 at step 210.
Original abstract
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.