One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions
AuthorsTomas Bruckner
Resources
A handful of one-token answers to simple prompts may be enough to tell whether an API is really serving the LLM it claims to use.
Key results
Number of models measured through OpenRouter.
Entropy in bits of single-token answer distributions.
Leave-one-out nearest-neighbor accuracy for documented model families.
Equal error rate using the 40-cell fingerprint battery.
Total responses collected for the ecosystem study.
What the paper found
Tomáš Bruckner presents a black-box method for fingerprinting and verifying large language models from the empirical distributions of single-token answers to trivial prompts such as “name a random number between 1 and 100.” Using Jensen–Shannon divergence across a 40-cell battery covering 10 tasks in English, Russian, Chinese, and Arabic, the study measures 165 models through OpenRouter, including OpenAI GPT-4o, Anthropic Claude Sonnet 5, Meta Llama, Qwen, and DeepSeek variants. The key insight is that models are systematically non-random in these tasks: their median cell entropy is 1.00 bit, and their answer distributions remain model-specific despite serving-provider variation, quantization, and temperature changes. Leave-one-out nearest-neighbor classification recovers documented model families with 59.5% accuracy, compared with an 18.4% chance rate. For identity verification, the full battery achieves an AUC of 0.971 and a 7.3% equal error rate; eight probe cells reduce the audit to roughly 120 single-token queries while producing a 10.6% equal error rate. The census collected 326,047 responses for $34.44, demonstrating that continuous endpoint auditing can be inexpensive. The analysis also detected ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model, showing that behavioral fingerprints can expose model substitution or unexpected serving configurations without logits, weights, long generations, adversarial prompts, or vendor cooperation.
Original abstract
Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights. Existing identification techniques require long generated texts, token-level log-probabilities, adversarially crafted prompts, or the model owner's cooperation. We show that far weaker evidence suffices. We define a behavioral fingerprint of an LLM as the empirical distribution of its answers to trivial one-word prompts - "name a random number between 1 and 100" - collected across four languages at a cost of one output token per query. Measuring 165 models served via a large commercial aggregator (OpenRouter), we find that (i) these distributions are highly non-uniform (median cell entropy 1.0 bit) and model-specific: split halves of the same model's samples lie an order of magnitude closer than samples of different models; (ii) Jensen-Shannon divergence between fingerprints recovers model lineage, assigning a model to its documented family with 59.5% leave-one-out accuracy against an 18.4% chance rate; and (iii) a biometric-style verification protocol achieves a 7.3% equal error rate with the full 40-cell battery, and below 11% with eight probe cells - roughly a hundred single-token queries per audit. We further report ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model. The protocol, prompts, raw data, and analysis code are released for reproduction and operational use.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.