NTH

PhantomBench: Benchmarking the Non-existential Threat of Language Models

AuthorsHaeji Jung, Hila Gonen

July 15, 2026 2 min read
Watch on YouTube
The one-line take

PhantomBench is a large benchmark that tests whether language models can recognize when a concept does not actually exist, revealing that even top models often confidently hallucinate.

Key results

62,411
PhantomBench concepts

Non-existent terms and entities in the benchmark

21
Evaluated models

Total language models evaluated across core and targeted experiments

86.7%
Peak hallucination rate

Highest reported average hallucination rate across evaluations

73.3%
Gemma 3 12B hallucination rate

Full-benchmark hallucination rate

7.3%
Llama 3.1 8B hallucination rate

Full-benchmark hallucination rate

0.755
Rare-concept correlation

Pearson correlation between abstention on non-existent and rare concepts

What the paper found

In PhantomBench, Haeji Jung and Hila Gonen of the University of British Columbia and Amii introduce a benchmark for testing whether language models recognize when they have no knowledge. The benchmark contains 62,411 linguistically plausible but non-existent terms and entities generated by blending words, recombining entity n-grams, and filtering exact matches against Dolma v1.7 with Infini-gram. The authors evaluate 21 models, including Llama, Gemma, Qwen, Mistral, Google’s Gemini 2.5, OpenAI’s GPT-OSS-20B, and DeepSeek-R1, using prompts about existence, meaning, date, and place. Hallucination rates reach 86.7%, and presupposing that a concept exists is especially damaging: models hallucinate far more when asked what an imaginary concept means than whether it exists. On the full benchmark, Gemma 3 12B has a 73.3% hallucination rate, while Llama 3.1 8B performs better at 7.3%, showing that model scale and capability do not guarantee calibrated abstention. Reasoning models and the largest variants in some families can hallucinate more, while domain-specialized models such as MedGemma and BioMistral are inconsistent. Most importantly, abstention on non-existent concepts correlates strongly with behavior on rare real concepts, with Pearson correlation 0.755, suggesting PhantomBench is a practical proxy for long-tail knowledge reliability in high-stakes applications.

Original abstract

Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them. This is particularly concerning in high-stakes domains, where consequences of such model behavior can lead to significant harms. Despite notable progress in understanding hallucinations, it remains unclear how reliably these models can recognize the limits of their knowledge. We introduce PhantomBench, the first large-scale benchmark of its kind, comprising more than 60K non-existent terms and entities derived from real concepts across diverse domains. Using our benchmark, we evaluate a total of 21 models of various types and sizes. We show staggering hallucination rates across the board (with average rates as high as 86.7% in some cases), and note that even frontier models surprisingly fail to abstain on non-existent concepts, especially when the input presumes their existence. We then show that PhantomBench can serve as a proxy for studying model behavior on rare concepts for which models are more prone to hallucinate. We also provide a pipeline to construct PhantomBench, enabling scalable generation of non-existent concepts tailored to the specific needs of researchers and practitioners.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis