Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
AuthorsWeifeng Yuan, Wenbo Guo, Feng Dong, Haoyu Wang, Yang Liu
Resources
LLM agents can invent nonexistent skills that attackers may pre-register as malicious packages, turning ordinary recommendation hallucinations into a supply-chain security threat.
Key results
Prompts used to measure hallucinated skill recommendations.
Standalone LLM and agent configurations tested.
Rate observed on real developer questions.
Unique fabricated skill names discovered.
Average repetitions of the dominant hallucinated name across 10 runs.
Average rate after retrieval grounding, down from 40.8%.
What the paper found
The paper presents the first large-scale study of skill-name hallucination in LLM agents: systems recommend downloadable skills whose names do not exist in any verified registry or GitHub repository. Evaluating 15,000 prompts across 12 configurations—including standalone models such as Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.4-mini, and GLM-4.5-Air, plus agents such as Claude Code, Codex, OpenClaw, and OpenCode—the authors find that every configuration hallucinates. On real developer questions, the rate reaches 43.1 percent, and the study identifies 5,669 distinct fabricated names. These are exploitable supply-chain targets: Claude Sonnet 4.6 repeated its dominant hallucinated name 7.8 times out of 10, while 410 names were shared across configurations and 15.0 percent of hallucinated names exactly matched PyPI or npm packages. Self-auditing achieved only 51 percent balanced accuracy, and hallucinated names were typically a median of six edits from real skills, defeating typosquatting filters. Four defenses expose a security-usability trade-off. Retrieval-augmented generation, using a real-skill index and ten retrieved candidates, reduced average hallucination from 40.8 percent to 3.2 percent, but mainly by suppressing outputs; even the strongest defended system recommended the correct skill no more than one time in six. Self-refinement was substantially weaker, while cross-model voting removed many hallucinations at the cost of rejecting valid skills. The authors conclude that prompt engineering and model scaling are insufficient, recommending registry-level name reservations and authenticated, exact-lookup recommendation pipelines for ecosystems developed by Anthropic, OpenAI, and related agent platforms.
Original abstract
LLM agents acquire new capabilities by downloading skills from open registries. Instead of browsing these catalogs manually, developers typically ask the agent to recommend and install a skill. This convenience hides a risk: agents frequently invent names for skills that exist in no registry. We term this flaw skill name hallucination. A fake name may seem harmless, but it opens the door to supply-chain attacks. Because registries rarely verify publishers, an adversary can prompt the agent, collect the fake names it returns, pre-register malicious skills under them, and wait for a victim to install the payload. We conducted the first large-scale measurement of skill name hallucination, evaluating 15,000 prompts across 12 configurations (4 standalone LLMs and 8 agents). We conservatively counted a name as hallucinated only if it was missing from all live registries and GitHub. The results reveal a systemic vulnerability: every configuration hallucinates. Rates average 36.0% for standalone LLMs and 36.9% for agents, rising to 43.1% on real-world developer questions. In total, the systems generated 5,669 distinct hallucinated names. Crucially, these names are not random noise. Agents repeat the same fake names across prompts and models, giving attackers highly reliable targets to hijack. Finally, we tested four model-level defenses and found a severe conflict between security and usability. The strongest, retrieval grounding, cut the hallucination rate from 40.8% to 3.2% but crippled usefulness: even the best-defended system recommended the correct skill only about one in six times. Skill name hallucination is thus a highly exploitable vulnerability requiring minimal attacker effort. Fixing it cannot rely on prompt engineering or model tuning alone. It demands ecosystem-wide structural changes: registry-level name reservations and verified recommendation pipelines.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.