Randomness in large language models: What researchers need to know (and report)
AuthorsGuillaume Coqueret, Joan Llull, Florian Oswald, Christophe Pérignon, Christoph Scheuch, Lars Vilhuber
Resources
This paper explains why LLM results can change unexpectedly and proposes practical rules for making LLM-based research more reproducible.
Key results
S&P 500 firms whose 10-K excerpts were classified
Identical prompt executions per model and filing experiment
Approximate share of GPT-5.6 regression t-statistics crossing the 90% threshold
Maximum accuracy difference in percentage points reported across repeated runs
Reported prime-versus-composite classification accuracy before model behavior changed
Reported accuracy three months later after a silent update
What the paper found
“Randomness in Large Language Models” argues that LLM-generated variables should be treated as draws from a distribution, not fixed measurements. Variation persists through deliberate sampling, silent provider updates, floating-point rounding affected by batching and hardware, model splitting, and mixture-of-experts routing. Setting temperature to zero removes deliberate sampling but not infrastructure or model changes; this distinction matters for proprietary APIs from OpenAI and Anthropic, where researchers cannot preserve the serving stack. The authors examine sentiment classification of 10-K filings from 100 S&P 500 firms, repeating the same prompt 200 times with OpenAI’s GPT-5.6 Luna, DeepSeek V4-Pro, and Google’s Gemma 4 26B A4B. With GPT-5.6, roughly 5% of regression t-statistics crossed a 90% significance threshold across repeated runs, showing how unstable LLM annotations can alter empirical conclusions. Consistent with prior evidence, nominally deterministic settings can produce accuracy differences of up to 15 percentage points, while GPT-4 performance on prime-number classification reportedly fell from 84% to 51% after a silent update. In the experiment, DeepSeek remained variable at temperature zero, whereas locally deployed Gemma became fully deterministic. The paper recommends archiving exact model versions, prompts, settings, timestamps, hardware and software environments, token and cost records, and every raw output. Open-weight systems such as Gemma, DeepSeek, and Meta’s Llama offer greater control than hosted models, but reproducibility still requires preserving the complete execution stack.
Original abstract
Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.