Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
AuthorsDifeng Ma, Changhua Pei, Yuanwei Lu, Quan Zhou, Zexin Wang, Yibo Zhu, Daxin Jiang, Dan Pei, Jingjing Li, Gaogang Xie
Resources
Instead of guessing when a GPU will fail, HeaRank identifies which GPUs are most likely to fail so operators can act before problems disrupt large-scale AI jobs.
Key results
These two failure types form the study’s primary target.
Maximum top-K F1 achieved by the evaluated prediction models.
Overall discrimination in production-scale host risk ranking.
HeaRank’s gain over LightGBM Ranker for the highest-priority nodes.
Share of future failures found within the top 5% of ranked nodes online.
Ranking performance under the longer labeling horizon.
What the paper found
Researchers from the Chinese Academy of Sciences, Tsinghua University, and StepFun argue that GPU reliability should be prioritized rather than precisely predicted. In production clusters, workload migration, non-stationary telemetry, low signal-to-noise ratios, and overlapping pre-failure and normal distributions make exact failure timing unreliable; even XGBoost, CNN, LSTM, Transformer, and Mixture-of-Experts models perform poorly, with the best top-K F1 reaching only 0.4837. The study focuses on Double Bit Errors and GPU Lost events, which account for 43.3% of service-impacting hardware incidents, and relates the problem to large-scale AI operations such as Meta’s LLaMA-3 training on 16,384 GPUs. The proposed HeaRank system reformulates reliability assessment as host-level risk ranking. Its compact MLP uses cumulative failure counts, recurrence and recency statistics, failure types, and hardware metadata instead of volatile time-series telemetry. On a production-scale, primarily NVIDIA cluster, HeaRank achieves an AUC of 0.834 and a 38% relative improvement in NDCG@5 over LightGBM Ranker. In six months of online deployment, it captured 64% of future failures within the top 5% of ranked nodes, compared with 21% for the incumbent health-score system. A 30-day labeling horizon raises AUC to 0.881, but the authors select a 7-day horizon because it better balances denoising with actionable maintenance scheduling. The central operational insight is to steer critical jobs toward lower-risk GPUs and reserve fragile nodes for tolerant workloads.
Original abstract
The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through an in-depth analysis of telemetry data from a productioncluster, we find that major GPU failures, including Double Bit Er-rors (DBEs) and GPU Lost events, exhibit strong stochasticity andlow signal-to-noise ratios in time-series telemetry, which makesconventional time-based prediction ineffective. This insight motivates a paradigm shift: instead of attempting topredict the absolute timing of a failure, we propose a more robustapproach focused on ranking nodes by their relative failure risk. Wepropose HeaRank (Health Rank), a Learning-to-Rank (LTR) frame-work that leverages stable historical failure patterns to computea global risk ranking of GPU nodes. Evaluated on a production-scale cluster with thousands of GPUs, HeaRank achieves an AUCof 0.83, significantly outperforming both heuristic baselines andstate-of-the-art ranking algorithms. In online deployment, HeaRanksuccessfully captures 64% of future failures within the top 5% ofranked nodes, compared to only 21% by the incumbent productionsystem. These results suggest that relative risk ranking can serveas a robust alternative in environments where absolute failure pre-diction is inherently limited. Our work highlights the importanceof risk-aware scheduling and proactive resource management inmodern GPU clusters.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.