NTH

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

AuthorsGaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li

AffiliationsUniversity of Tennessee, Knoxville · University of Chicago · Amazon · University of Texas at Austin

September 26, 2026 3 min read
Watch on YouTube
The one-line take

Even supposedly deterministic greedy LLM decoding can produce different answers in BF16 and FP16, but selectively recomputing the final layer in FP32 can greatly improve agreement at low cost.

Key results

49%
Minimum divergence rate

Lower bound of BF16-versus-FP16 prompt divergence across the evaluated models and benchmarks.

41%
TinyLlama GSM8K baseline agreement

Exact sequence agreement between BF16 and FP16 greedy decoding.

22
TinyLlama FP32 lm_head gain

Percentage-point improvement in GSM8K exact agreement from selective FP32 lm_head recomputation.

36
Llama-3.2-3B FP32 lm_head gain

Percentage-point cross-precision agreement improvement on GSM8K.

4%
Selective recomputation overhead

Latency overhead remains below this level in low-batch NVIDIA A10G inference.

8
Batch-size failure boundary

At batch size 8, the lm_head-only intervention provides zero lift in the tested setup.

What the paper found

The paper demonstrates that greedy decoding is deterministic only within a fixed numerical format: switching the same LLM, prompt, hardware, and decoding settings from BF16 to FP16 can change the output. Across six models from the Llama, Qwen, Mistral, and OLMoE families, including DeepSeek-R1-Distill-Qwen-7B, at least 49% of prompts diverged, and a single token flip often cascaded into a different autoregressive trajectory. On TinyLlama-1.1B, exact agreement was just 41% on GSM8K, 36% on HumanEval, and 30% on MBPP. The mechanism is not general transformer-body error: divergence concentrates at the final lm_head when the top-two logit margin is small relative to precision-dependent perturbations, with the measured margin separating divergent and agreed cases by 150 times. The proposed remedy selectively recomputes the lm_head in FP32 only when the margin falls below 10^-3; on TinyLlama, this raises GSM8K agreement by 22 percentage points, while the broader cross-model gain reaches 36 percentage points on Llama-3.2-3B, with under 4% latency overhead on an NVIDIA A10G. Top-K FP32 recomputation with K=2 matches full-vocabulary recomputation, showing that correcting the leading candidates is sufficient in the tested regime. The method is partial: its benefit falls to zero at batch size 8 and on body-dominated cases such as end-to-end FP8 or BF16-saturated Qwen variants. For deployment, the study recommends recording precision and hardware explicitly rather than assuming that greedy outputs from systems using models such as Llama, Qwen, Mistral, or DeepSeek will replay identically.

Original abstract

Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis