Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
AuthorsHaoyaun Zhu, Jie Zhang
Resources
LLM judges may not be stable measuring tools at all, because identical requests can produce different rankings across time, providers, and server conditions.
Key results
Spearman median, versus the preregistered 0.90 stability threshold
Exact full-ranking matches out of 100 byte-identical replays, versus a 0.99 threshold
Maximum simulated observer calls; the gate passed 0 of 500 replicates
Increase under concurrent load on the self-hosted vLLM 0.11.0 arm
What the paper found
This preregistered audit tested whether black-box LLM judges behave like stable scientific instruments when the same request is sent repeatedly to a shared endpoint. Across 52988 audited request attempts, the engineering pipeline was effectively perfect—requests, schemas, hashes, and metadata were valid—yet the measurement gates failed: same-window repeat rankings reached only Spearman 0.400 against a 0.90 requirement, while byte-identical next-day replays matched exactly 0.78 of the time against a 0.99 requirement. The authors identify three interacting causes: label-to-meaning bias in a two-label log-probability readout, candidate score gaps far smaller than the observer’s noise floor, and endpoint nondeterminism amplified by exact-permutation rankings. A 748000-call simulation passed the ranking gate 0 of 500 times, showing that more sampling could not repair a near-degenerate measurand. Follow-up tests found the same stability floor across OpenAI gpt-4.1, DeepSeek, Mistral, and Qwen endpoints, while provider metadata—including fingerprints—did not predict replay agreement. Self-hosting with vLLM 0.11.0 improved quiet-server agreement, but concurrent load increased disagreement 8.4-fold. The practical contribution is an instrument-first protocol: pilot noise floors and score gaps, verify snapshot identity and metadata semantics, calibrate operating characteristics, and use continuous or aggregated readouts before freezing evaluation gates.
Original abstract
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.