NTH

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

AuthorsAndrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, Seth Donoughe

June 13, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces ABC-Bench, a benchmark that tests whether LLM agents can carry out biology tasks with biosecurity implications, and shows they can already perform surprisingly well, even in some wet-lab validated cases.

Key results

0.33
human baseline Fragment Design

mean expert baseliner score

0.22
human baseline Screening Evasion

mean expert baseliner score

0.20
human baseline Liquid Handling Robot

mean expert baseliner score

175
baselining effort

total expert person-hours

1.00
best task score

perfect scores achieved on benchmark tasks by frontier models such as Claude Opus 4.6, Claude Sonnet 4.6, and Gemini 3.1 Pro Preview

What the paper found

ABC-Bench, from SecureBio, is an agentic biosecurity benchmark that tests frontier AI models on three tool-using biology tasks: Fragment Design, Screening Evasion, and a Liquid Handling Robot protocol for Gibson Assembly on an OpenTrons OT-2. The benchmark is built around a risk chain from fragment design to synthesis-screening evasion to wet-lab execution, and it emphasizes objective scoring through automated checks plus real-world validation. In evaluation, the paper reports that all tested models outperformed the median human baseliner across the three tasks, with baseline human mean scores of 0.33 on Fragment Design, 0.22 on Screening Evasion, and 0.20 on Liquid Handling Robot, after 175 person-hours of expert baselining. Claude Sonnet 4.6 and Gemini 3.1 Pro Preview reached perfect Liquid Handling Robot scores of 1.00, Claude Opus 4.6 achieved 1.00 on Fragment Design, and Gemini 3.1 Pro Preview reached 0.78 on Screening Evasion. The most striking safety-relevant result is that GPT-o4-mini-high generated OpenTrons scripts that successfully performed DNA assembly in three independent wet-lab experiments, confirmed by whole-plasmid sequencing. The benchmark also exposes refusal behavior: on Screening Evasion, Claude Sonnet 4.6, Claude Opus 4.6, and GPT-5.4 refused all samples, showing that some frontier systems recognized the dual-use nature of the task while still demonstrating strong bio-capability on the other tasks.

Original abstract

Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biology tasks that previously required experienced human biologists. These emerging AI capabilities offer new opportunities for scientific discovery and biomedical advances, but they also shift the landscape of biosecurity risks. To address this, we introduce the Agentic Bio-Capabilities Benchmark (ABC-Bench), a suite of tasks to measure agentic biosecurity-relevant capabilities. ABC-Bench evaluates LLM agents on both benign and dual-use biology tasks: writing code to operate liquid handling robots, designing DNA fragments for in vitro assembly, and evading DNA synthesis screening. These tasks require a combination of biology and software expertise. All tested LLM agents outperformed the median expert human baseliner on all three tasks. Agents performed highly on tasks drawing on published knowledge and well-documented protocols, and more weakly on a task requiring novel bioinformatics reasoning. In three wet-lab validation experiments, we found that OpenAI's o4-mini-high produced scripts that, when run on an OpenTrons liquid handling robot, successfully assembled DNA with expected sequences.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis