Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
AuthorsQian Zhao, Kunlong Chen, Changxin Tian, Zhonghui Jiang, Haitao Zhang, Chaofan Yu, Peijie Jiang, Mingliang Gong, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou
Resources
This paper explains why some 4-bit LLM training formats quietly introduce harmful rounding bias, and shows that a more uniform 4-bit recipe can make pretraining more stable and efficient.
Key results
E2M1 baseline latest-1000-step loss error
UFP4 latest-1000-step loss error
E2M1 baseline latest-1000-step loss error
UFP4 latest-1000-step loss error
E2M1 baseline latest-1000-step loss error
UFP4 latest-1000-step loss error
What the paper found
This Ant Group Ling Team paper argues that the default FP4 training choice in NVIDIA Blackwell/Rubin-class and AMD MI350-style recipes, E2M1, has a geometric flaw: its non-uniform RTNE bins induce “Shrinkage Bias,” a systematic negative rounding error that compounds multiplicatively across layers. The authors show that Random Hadamard Transform, which is used to improve bucket utilization, can actually worsen E2M1 because it pushes tensor mass into E2M1’s most asymmetric bins, turning outlier-mitigation into local-resolution failure. They then introduce UFP4, a uniform 4-bit recipe built on E1M2/INT4-style grids, applying RHT to all three training GEMMs—FPROP, DGRAD, and WGRAD—while keeping stochastic rounding only on dY. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 stays closer to BF16, reducing latest-1000-step BF16-relative loss error from 1.2570% to 0.9673%, from 2.3596% to 1.8469%, and from 1.7308% to 1.3863%, respectively. An ablation on Dense 1.5B shows full-RHT plus SR is best, improving mean LM loss by 0.01123 over no RHT, and fused RHT-plus-quantization adds only about 1.06x to 1.07x latency over standalone quantization. The paper’s central claim is that future accelerators should treat E1M2/INT4-style uniform 4-bit grids as first-class training primitives, not just E2M1.
Original abstract
FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geometric asymmetry of their representable bins. We show that this bias accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform (RHT), providing a unified explanation for the training instability observed in existing E2M1-based FP4 recipes. In contrast, uniform grids (E1M2/INT4) bypass this grid-geometry error and better convert the improved bucket utilization from RHT into higher quantization quality. Based on this finding, we propose UFP4, a uniform 4-bit training recipe that applies RHT to all three training GEMMs while restricting stochastic rounding to dY alone. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 consistently achieves lower BF16-relative loss degradation than strong E2M1-based baselines, supported by scaling-law analysis and ablation studies. Our results suggest that future accelerators should support E1M2/INT4-style uniform 4-bit grids as first-class training primitives alongside E2M1.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.