NTH
Research collection

AI Safety research

Research on model reliability, alignment, misuse, and robustness. Examine evaluation methods and the evidence behind proposed safeguards.

39 papers · Latest edition October 7, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All AI Safety papers

Newest editions first.

02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis
06Safety

MOLE: Detecting Insider Threats in AI Agents

Aashiq Muhamed, Virginia Smith

MOLE tests whether security monitors can catch AI agents secretly stealing data or sabotaging systems—and finds that even strong monitors miss much of the harm.

Read analysis
11Safety

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Xiaodong Wu, Yu Shi, Qi Li, Zhimin Zhao, Xiangman Li, Bram Adams, Ahmed E. Hassan, Jianbing Ni

EVOMAL shows how coding agents can unknowingly copy and spread malicious skills through their own self-improvement loops, while proposing a prompt-based defense to contain the attack.

Read analysis
13Safety

Inadvertent Context Leakage in Language Models

Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri, Sanjam Garg, Saeed Mahloujifar

The study shows that language models may leak secrets indirectly through ordinary responses, even when they refuse to reveal them explicitly.

Read analysis
17Safety

Large language models improve physician accuracy but lead to false reliance

Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker

Citation-backed medical LLMs can substantially improve physician accuracy, but they may also make doctors less likely to challenge confidently supported wrong answers.

Read analysis
22Safety

AgentAbstain: Do LLM Agents Know When Not to Act?

Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran

AgentAbstain tests whether LLM agents can recognize when they should not act, revealing that even strong agents often take risky actions when abstention is required.

Read analysis
25Safety

LLM Evaluators are Biased across Languages

Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen

Multilingual LLM judges can score identical content very differently across languages, allowing harmful material in lower-resource languages to slip past safety filters despite impressive benchmark accuracy.

Read analysis
27Safety

Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents

A H M Nazmus Sakib, Dipayan Banik, Murtuza Jadliwala

This study finds that autonomous coding workflows frequently contain overlooked security problems, especially supply-chain issues and leaked credentials, calling for stronger safeguards at the human-AI collaboration boundary.

Read analysis
29Safety

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa, Christian Schroeder de Witt

This paper shows that LLM agents with tools can secretly coordinate through steganographic channels that evade monitoring, raising serious new safety risks for multi-agent systems.

Read analysis
30Safety

Agent Data Injection Attacks are Realistic Threats to AI Agents

Woohyuk Choi, Juhee Kim, Taehyun Kang, Jihyeon Jeong, Luyi Xing, Byoungyoung Lee

This paper shows that AI agents can be tricked not just by malicious instructions, but by seemingly trustworthy data that secretly steers them into unsafe actions, exposing real vulnerabilities in popular agents.

Read analysis
31Safety

GPUBreach: Privilege Escalation Attacks on GPUs using Rowhammer

Chris S. Lin, Yuqin Yan, Guozhen Ding, Joyce Qu, Joseph Zhu, David Lie, Gururaj Saileshwar

This paper shows that GPU Rowhammer can do more than break models: it can let an attacker steal data, tamper with GPU code, and even potentially break out to root on the host machine.

Read analysis
35Safety

AI systems out-persuade expert humans

Kobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin, Ed Saunders, Hannah Rose Kirk, Helen Margetts, Christopher Summerfield

This study shows that AI can already out-persuade even expert humans in conversation, raising major questions about how persuasive systems could be used in politics, fundraising, and everyday influence.

Read analysis
36Safety

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, Seth Donoughe

This paper introduces ABC-Bench, a benchmark that tests whether LLM agents can carry out biology tasks with biosecurity implications, and shows they can already perform surprisingly well, even in some wet-lab validated cases.

Read analysis
37Safety

AI Agents Enable Adaptive Computer Worms

Jonas Guan, Tom Blanchard, Hanna Foerster, Hengrui Jia, Gabriel Huang, Nicolas Papernot

This paper argues that AI agents can power a new kind of self-sustaining worm that adapts to each target, spreads across real networks, and uses stolen compute to keep attacking at near-zero marginal cost.

Read analysis