NTH

Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents

AuthorsWeifeng Yuan, Wenbo Guo, Feng Dong, Haoyu Wang, Yang Liu

July 17, 2026 3 min read
Watch on YouTube
The one-line take

LLM agents can invent nonexistent skills that attackers may pre-register as malicious packages, turning ordinary recommendation hallucinations into a supply-chain security threat.

Key results

15,000
Evaluation prompts

Prompts used to measure hallucinated skill recommendations.

12
Evaluated configurations

Standalone LLM and agent configurations tested.

43.1%
Real-world hallucination rate

Rate observed on real developer questions.

5,669
Distinct hallucinated names

Unique fabricated skill names discovered.

7.8
Claude Sonnet recurrence

Average repetitions of the dominant hallucinated name across 10 runs.

3.2%
RAG hallucination rate

Average rate after retrieval grounding, down from 40.8%.

What the paper found

The paper presents the first large-scale study of skill-name hallucination in LLM agents: systems recommend downloadable skills whose names do not exist in any verified registry or GitHub repository. Evaluating 15,000 prompts across 12 configurations—including standalone models such as Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.4-mini, and GLM-4.5-Air, plus agents such as Claude Code, Codex, OpenClaw, and OpenCode—the authors find that every configuration hallucinates. On real developer questions, the rate reaches 43.1 percent, and the study identifies 5,669 distinct fabricated names. These are exploitable supply-chain targets: Claude Sonnet 4.6 repeated its dominant hallucinated name 7.8 times out of 10, while 410 names were shared across configurations and 15.0 percent of hallucinated names exactly matched PyPI or npm packages. Self-auditing achieved only 51 percent balanced accuracy, and hallucinated names were typically a median of six edits from real skills, defeating typosquatting filters. Four defenses expose a security-usability trade-off. Retrieval-augmented generation, using a real-skill index and ten retrieved candidates, reduced average hallucination from 40.8 percent to 3.2 percent, but mainly by suppressing outputs; even the strongest defended system recommended the correct skill no more than one time in six. Self-refinement was substantially weaker, while cross-model voting removed many hallucinations at the cost of rejecting valid skills. The authors conclude that prompt engineering and model scaling are insufficient, recommending registry-level name reservations and authenticated, exact-lookup recommendation pipelines for ecosystems developed by Anthropic, OpenAI, and related agent platforms.

Original abstract

LLM agents acquire new capabilities by downloading skills from open registries. Instead of browsing these catalogs manually, developers typically ask the agent to recommend and install a skill. This convenience hides a risk: agents frequently invent names for skills that exist in no registry. We term this flaw skill name hallucination. A fake name may seem harmless, but it opens the door to supply-chain attacks. Because registries rarely verify publishers, an adversary can prompt the agent, collect the fake names it returns, pre-register malicious skills under them, and wait for a victim to install the payload. We conducted the first large-scale measurement of skill name hallucination, evaluating 15,000 prompts across 12 configurations (4 standalone LLMs and 8 agents). We conservatively counted a name as hallucinated only if it was missing from all live registries and GitHub. The results reveal a systemic vulnerability: every configuration hallucinates. Rates average 36.0% for standalone LLMs and 36.9% for agents, rising to 43.1% on real-world developer questions. In total, the systems generated 5,669 distinct hallucinated names. Crucially, these names are not random noise. Agents repeat the same fake names across prompts and models, giving attackers highly reliable targets to hijack. Finally, we tested four model-level defenses and found a severe conflict between security and usability. The strongest, retrieval grounding, cut the hallucination rate from 40.8% to 3.2% but crippled usefulness: even the best-defended system recommended the correct skill only about one in six times. Skill name hallucination is thus a highly exploitable vulnerability requiring minimal attacker effort. Fixing it cannot rely on prompt engineering or model tuning alone. It demands ecosystem-wide structural changes: registry-level name reservations and verified recommendation pipelines.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis