NTH

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

AuthorsQian Zhao, Kunlong Chen, Changxin Tian, Zhonghui Jiang, Haitao Zhang, Chaofan Yu, Peijie Jiang, Mingliang Gong, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

June 25, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains why some 4-bit LLM training formats quietly introduce harmful rounding bias, and shows that a more uniform 4-bit recipe can make pretraining more stable and efficient.

Key results

1.2570%
Dense 1.5B BF16-relative error

E2M1 baseline latest-1000-step loss error

0.9673%
Dense 1.5B BF16-relative error

UFP4 latest-1000-step loss error

2.3596%
MoE 7.9B BF16-relative error

E2M1 baseline latest-1000-step loss error

1.8469%
MoE 7.9B BF16-relative error

UFP4 latest-1000-step loss error

1.7308%
MoE 124B BF16-relative error

E2M1 baseline latest-1000-step loss error

1.3863%
MoE 124B BF16-relative error

UFP4 latest-1000-step loss error

What the paper found

This Ant Group Ling Team paper argues that the default FP4 training choice in NVIDIA Blackwell/Rubin-class and AMD MI350-style recipes, E2M1, has a geometric flaw: its non-uniform RTNE bins induce “Shrinkage Bias,” a systematic negative rounding error that compounds multiplicatively across layers. The authors show that Random Hadamard Transform, which is used to improve bucket utilization, can actually worsen E2M1 because it pushes tensor mass into E2M1’s most asymmetric bins, turning outlier-mitigation into local-resolution failure. They then introduce UFP4, a uniform 4-bit recipe built on E1M2/INT4-style grids, applying RHT to all three training GEMMs—FPROP, DGRAD, and WGRAD—while keeping stochastic rounding only on dY. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 stays closer to BF16, reducing latest-1000-step BF16-relative loss error from 1.2570% to 0.9673%, from 2.3596% to 1.8469%, and from 1.7308% to 1.3863%, respectively. An ablation on Dense 1.5B shows full-RHT plus SR is best, improving mean LM loss by 0.01123 over no RHT, and fused RHT-plus-quantization adds only about 1.06x to 1.07x latency over standalone quantization. The paper’s central claim is that future accelerators should treat E1M2/INT4-style uniform 4-bit grids as first-class training primitives, not just E2M1.

Original abstract

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geometric asymmetry of their representable bins. We show that this bias accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform (RHT), providing a unified explanation for the training instability observed in existing E2M1-based FP4 recipes. In contrast, uniform grids (E1M2/INT4) bypass this grid-geometry error and better convert the improved bucket utilization from RHT into higher quantization quality. Based on this finding, we propose UFP4, a uniform 4-bit training recipe that applies RHT to all three training GEMMs while restricting stochastic rounding to dY alone. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 consistently achieves lower BF16-relative loss degradation than strong E2M1-based baselines, supported by scaling-law analysis and ablation studies. Our results suggest that future accelerators should support E1M2/INT4-style uniform 4-bit grids as first-class training primitives alongside E2M1.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis