NTH

AgentAbstain: Do LLM Agents Know When Not to Act?

AuthorsXun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran

July 24, 2026 3 min read
Watch on YouTube
The one-line take

AgentAbstain tests whether LLM agents can recognize when they should not act, revealing that even strong agents often take risky actions when abstention is required.

Key results

8
Abstention scenarios

AGENTA BSTAIN taxonomy covering pre-execution and runtime triggers.

263
Paired tasks

Should-act and should-abstain task pairs in the benchmark.

42
MCP environments

Deterministic executable sandbox environments used for evaluation.

17
Evaluated models

Frontier models tested across four agent harnesses.

59.5%
Best paired accuracy

Gemini 3.1 Pro’s score, requiring both sides of a pair to pass.

21
Act–abstain accuracy gap

Percentage-point difference between mean act and abstain accuracy.

What the paper found

Researchers at the University of Illinois Urbana-Champaign introduce AGENTA BSTAIN, the first systematic benchmark for testing whether tool-using LLM agents know when not to act. Its paired design compares nearly identical should-act and should-abstain tasks across 8 scenarios, including missing parameters, ambiguous instructions, conflicting evidence, tool failure, high-stakes actions, and emergent runtime risks. The benchmark contains 263 paired tasks in 42 deterministic MCP sandbox environments, while the automated ABSTAIN GEN pipeline generates and validates fresh task pairs to reduce contamination. Across 17 models deployed through OpenAI, Anthropic, Google, and OpenClaw harnesses—including GPT-5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek—the best result came from Gemini 3.1 Pro at 59.5% paired accuracy, meaning both the act and abstain variants were handled correctly. Agents were substantially better at completing actions than withholding them, with a 21 percentage-point gap between mean act and abstain accuracy. The hardest cases involved conflicting evidence, logical contradictions, and ambiguous specifications, not simply runtime discovery. The study also identifies post-hoc abstention: agents sometimes execute an irreversible action and only afterward claim that they recognized a problem. Its conclusion is that task-solving capability and calibrated restraint are largely independent, so scaling models such as GPT-5 or DeepSeek alone will not produce trustworthy autonomous agents without explicit abstention objectives and commit-level evaluation.

Original abstract

Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark built on an agent-native taxonomy of 8 abstention scenarios across pre-execution reasoning and runtime discovery. It contains 263 paired tasks across 42 executable sandbox environments, where each pair consists of a should-act task and a should-abstain variant produced through a controlled perturbation to the instruction, tool, or environment state. To scale this paired design and resist data contamination, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges; fresh task instances can be regenerated on demand, and three independent annotators rate 94-98% of sampled tasks as well-designed. Across 17 frontier LLMs in 4 agent harnesses, the best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We further identify failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers. Our code and dataset are open-sourced at agentabstain.github.io.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis