NTH

One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

AuthorsTomas Bruckner

July 24, 2026 3 min read
Watch on YouTube
The one-line take

A handful of one-token answers to simple prompts may be enough to tell whether an API is really serving the LLM it claims to use.

Key results

165
Models surveyed

Number of models measured through OpenRouter.

1.00
Median cell entropy

Entropy in bits of single-token answer distributions.

59.5%
Family classification accuracy

Leave-one-out nearest-neighbor accuracy for documented model families.

7.3%
Full-battery verification EER

Equal error rate using the 40-cell fingerprint battery.

326047
Census responses

Total responses collected for the ecosystem study.

What the paper found

Tomáš Bruckner presents a black-box method for fingerprinting and verifying large language models from the empirical distributions of single-token answers to trivial prompts such as “name a random number between 1 and 100.” Using Jensen–Shannon divergence across a 40-cell battery covering 10 tasks in English, Russian, Chinese, and Arabic, the study measures 165 models through OpenRouter, including OpenAI GPT-4o, Anthropic Claude Sonnet 5, Meta Llama, Qwen, and DeepSeek variants. The key insight is that models are systematically non-random in these tasks: their median cell entropy is 1.00 bit, and their answer distributions remain model-specific despite serving-provider variation, quantization, and temperature changes. Leave-one-out nearest-neighbor classification recovers documented model families with 59.5% accuracy, compared with an 18.4% chance rate. For identity verification, the full battery achieves an AUC of 0.971 and a 7.3% equal error rate; eight probe cells reduce the audit to roughly 120 single-token queries while producing a 10.6% equal error rate. The census collected 326,047 responses for $34.44, demonstrating that continuous endpoint auditing can be inexpensive. The analysis also detected ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model, showing that behavioral fingerprints can expose model substitution or unexpected serving configurations without logits, weights, long generations, adversarial prompts, or vendor cooperation.

Original abstract

Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights. Existing identification techniques require long generated texts, token-level log-probabilities, adversarially crafted prompts, or the model owner's cooperation. We show that far weaker evidence suffices. We define a behavioral fingerprint of an LLM as the empirical distribution of its answers to trivial one-word prompts - "name a random number between 1 and 100" - collected across four languages at a cost of one output token per query. Measuring 165 models served via a large commercial aggregator (OpenRouter), we find that (i) these distributions are highly non-uniform (median cell entropy 1.0 bit) and model-specific: split halves of the same model's samples lie an order of magnitude closer than samples of different models; (ii) Jensen-Shannon divergence between fingerprints recovers model lineage, assigning a model to its documented family with 59.5% leave-one-out accuracy against an 18.4% chance rate; and (iii) a biometric-style verification protocol achieves a 7.3% equal error rate with the full 40-cell battery, and below 11% with eight probe cells - roughly a hundred single-token queries per audit. We further report ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model. The protocol, prompts, raw data, and analysis code are released for reproduction and operational use.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis