ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
AuthorsAndrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, Seth Donoughe
Resources
This paper introduces ABC-Bench, a benchmark that tests whether LLM agents can carry out biology tasks with biosecurity implications, and shows they can already perform surprisingly well, even in some wet-lab validated cases.
Key results
mean expert baseliner score
mean expert baseliner score
mean expert baseliner score
total expert person-hours
perfect scores achieved on benchmark tasks by frontier models such as Claude Opus 4.6, Claude Sonnet 4.6, and Gemini 3.1 Pro Preview
What the paper found
ABC-Bench, from SecureBio, is an agentic biosecurity benchmark that tests frontier AI models on three tool-using biology tasks: Fragment Design, Screening Evasion, and a Liquid Handling Robot protocol for Gibson Assembly on an OpenTrons OT-2. The benchmark is built around a risk chain from fragment design to synthesis-screening evasion to wet-lab execution, and it emphasizes objective scoring through automated checks plus real-world validation. In evaluation, the paper reports that all tested models outperformed the median human baseliner across the three tasks, with baseline human mean scores of 0.33 on Fragment Design, 0.22 on Screening Evasion, and 0.20 on Liquid Handling Robot, after 175 person-hours of expert baselining. Claude Sonnet 4.6 and Gemini 3.1 Pro Preview reached perfect Liquid Handling Robot scores of 1.00, Claude Opus 4.6 achieved 1.00 on Fragment Design, and Gemini 3.1 Pro Preview reached 0.78 on Screening Evasion. The most striking safety-relevant result is that GPT-o4-mini-high generated OpenTrons scripts that successfully performed DNA assembly in three independent wet-lab experiments, confirmed by whole-plasmid sequencing. The benchmark also exposes refusal behavior: on Screening Evasion, Claude Sonnet 4.6, Claude Opus 4.6, and GPT-5.4 refused all samples, showing that some frontier systems recognized the dual-use nature of the task while still demonstrating strong bio-capability on the other tasks.
Original abstract
Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biology tasks that previously required experienced human biologists. These emerging AI capabilities offer new opportunities for scientific discovery and biomedical advances, but they also shift the landscape of biosecurity risks. To address this, we introduce the Agentic Bio-Capabilities Benchmark (ABC-Bench), a suite of tasks to measure agentic biosecurity-relevant capabilities. ABC-Bench evaluates LLM agents on both benign and dual-use biology tasks: writing code to operate liquid handling robots, designing DNA fragments for in vitro assembly, and evading DNA synthesis screening. These tasks require a combination of biology and software expertise. All tested LLM agents outperformed the median expert human baseliner on all three tasks. Agents performed highly on tasks drawing on published knowledge and well-documented protocols, and more weakly on a task requiring novel bioinformatics reasoning. In three wet-lab validation experiments, we found that OpenAI's o4-mini-high produced scripts that, when run on an OpenTrons liquid handling robot, successfully assembled DNA with expected sequences.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.