PhantomBench: Benchmarking the Non-existential Threat of Language Models
AuthorsHaeji Jung, Hila Gonen
Resources
PhantomBench is a large benchmark that tests whether language models can recognize when a concept does not actually exist, revealing that even top models often confidently hallucinate.
Key results
Non-existent terms and entities in the benchmark
Total language models evaluated across core and targeted experiments
Highest reported average hallucination rate across evaluations
Full-benchmark hallucination rate
Full-benchmark hallucination rate
Pearson correlation between abstention on non-existent and rare concepts
What the paper found
In PhantomBench, Haeji Jung and Hila Gonen of the University of British Columbia and Amii introduce a benchmark for testing whether language models recognize when they have no knowledge. The benchmark contains 62,411 linguistically plausible but non-existent terms and entities generated by blending words, recombining entity n-grams, and filtering exact matches against Dolma v1.7 with Infini-gram. The authors evaluate 21 models, including Llama, Gemma, Qwen, Mistral, Google’s Gemini 2.5, OpenAI’s GPT-OSS-20B, and DeepSeek-R1, using prompts about existence, meaning, date, and place. Hallucination rates reach 86.7%, and presupposing that a concept exists is especially damaging: models hallucinate far more when asked what an imaginary concept means than whether it exists. On the full benchmark, Gemma 3 12B has a 73.3% hallucination rate, while Llama 3.1 8B performs better at 7.3%, showing that model scale and capability do not guarantee calibrated abstention. Reasoning models and the largest variants in some families can hallucinate more, while domain-specialized models such as MedGemma and BioMistral are inconsistent. Most importantly, abstention on non-existent concepts correlates strongly with behavior on rare real concepts, with Pearson correlation 0.755, suggesting PhantomBench is a practical proxy for long-tail knowledge reliability in high-stakes applications.
Original abstract
Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them. This is particularly concerning in high-stakes domains, where consequences of such model behavior can lead to significant harms. Despite notable progress in understanding hallucinations, it remains unclear how reliably these models can recognize the limits of their knowledge. We introduce PhantomBench, the first large-scale benchmark of its kind, comprising more than 60K non-existent terms and entities derived from real concepts across diverse domains. Using our benchmark, we evaluate a total of 21 models of various types and sizes. We show staggering hallucination rates across the board (with average rates as high as 86.7% in some cases), and note that even frontier models surprisingly fail to abstain on non-existent concepts, especially when the input presumes their existence. We then show that PhantomBench can serve as a proxy for studying model behavior on rare concepts for which models are more prone to hallucinate. We also provide a pipeline to construct PhantomBench, enabling scalable generation of non-existent concepts tailored to the specific needs of researchers and practitioners.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.