NTH

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

AuthorsHaoyaun Zhu, Jie Zhang

September 7, 2026 3 min read
Watch on YouTube
The one-line take

LLM judges may not be stable measuring tools at all, because identical requests can produce different rankings across time, providers, and server conditions.

Key results

0.400
Same-window ranking agreement

Spearman median, versus the preregistered 0.90 stability threshold

0.78
Next-day exact replay agreement

Exact full-ranking matches out of 100 byte-identical replays, versus a 0.99 threshold

748000
Sampling simulation

Maximum simulated observer calls; the gate passed 0 of 500 replicates

8.4-fold
Load-induced disagreement increase

Increase under concurrent load on the self-hosted vLLM 0.11.0 arm

What the paper found

This preregistered audit tested whether black-box LLM judges behave like stable scientific instruments when the same request is sent repeatedly to a shared endpoint. Across 52988 audited request attempts, the engineering pipeline was effectively perfect—requests, schemas, hashes, and metadata were valid—yet the measurement gates failed: same-window repeat rankings reached only Spearman 0.400 against a 0.90 requirement, while byte-identical next-day replays matched exactly 0.78 of the time against a 0.99 requirement. The authors identify three interacting causes: label-to-meaning bias in a two-label log-probability readout, candidate score gaps far smaller than the observer’s noise floor, and endpoint nondeterminism amplified by exact-permutation rankings. A 748000-call simulation passed the ranking gate 0 of 500 times, showing that more sampling could not repair a near-degenerate measurand. Follow-up tests found the same stability floor across OpenAI gpt-4.1, DeepSeek, Mistral, and Qwen endpoints, while provider metadata—including fingerprints—did not predict replay agreement. Self-hosting with vLLM 0.11.0 improved quiet-server agreement, but concurrent load increased disagreement 8.4-fold. The practical contribution is an instrument-first protocol: pilot noise floors and score gaps, verify snapshot identity and metadata semantics, calibrate operating characteristics, and use continuous or aggregated readouts before freezing evaluation gates.

Original abstract

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis