NTH

Randomness in large language models: What researchers need to know (and report)

AuthorsGuillaume Coqueret, Joan Llull, Florian Oswald, Christophe Pérignon, Christoph Scheuch, Lars Vilhuber

July 31, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains why LLM results can change unexpectedly and proposes practical rules for making LLM-based research more reproducible.

Key results

100
Filing sample

S&P 500 firms whose 10-K excerpts were classified

200
Repeated runs

Identical prompt executions per model and filing experiment

5%
Significance-threshold crossings

Approximate share of GPT-5.6 regression t-statistics crossing the 90% threshold

15
Deterministic-setting variation

Maximum accuracy difference in percentage points reported across repeated runs

84%
GPT-4 initial accuracy

Reported prime-versus-composite classification accuracy before model behavior changed

51%
GPT-4 later accuracy

Reported accuracy three months later after a silent update

What the paper found

“Randomness in Large Language Models” argues that LLM-generated variables should be treated as draws from a distribution, not fixed measurements. Variation persists through deliberate sampling, silent provider updates, floating-point rounding affected by batching and hardware, model splitting, and mixture-of-experts routing. Setting temperature to zero removes deliberate sampling but not infrastructure or model changes; this distinction matters for proprietary APIs from OpenAI and Anthropic, where researchers cannot preserve the serving stack. The authors examine sentiment classification of 10-K filings from 100 S&P 500 firms, repeating the same prompt 200 times with OpenAI’s GPT-5.6 Luna, DeepSeek V4-Pro, and Google’s Gemma 4 26B A4B. With GPT-5.6, roughly 5% of regression t-statistics crossed a 90% significance threshold across repeated runs, showing how unstable LLM annotations can alter empirical conclusions. Consistent with prior evidence, nominally deterministic settings can produce accuracy differences of up to 15 percentage points, while GPT-4 performance on prime-number classification reportedly fell from 84% to 51% after a silent update. In the experiment, DeepSeek remained variable at temperature zero, whereas locally deployed Gemma became fully deterministic. The paper recommends archiving exact model versions, prompts, settings, timestamps, hardware and software environments, token and cost records, and every raw output. Open-weight systems such as Gemma, DeepSeek, and Meta’s Llama offer greater control than hosted models, but reproducibility still requires preserving the complete execution stack.

Original abstract

Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis