NTH

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

AuthorsDifeng Ma, Changhua Pei, Yuanwei Lu, Quan Zhou, Zexin Wang, Yibo Zhu, Daxin Jiang, Dan Pei, Jingjing Li, Gaogang Xie

July 24, 2026 3 min read
Watch on YouTube
The one-line take

Instead of guessing when a GPU will fail, HeaRank identifies which GPUs are most likely to fail so operators can act before problems disrupt large-scale AI jobs.

Key results

43.3%
Critical incident share from DBE and GPU Lost

These two failure types form the study’s primary target.

0.4837
Best time-series prediction F1

Maximum top-K F1 achieved by the evaluated prediction models.

0.834
HeaRank AUC

Overall discrimination in production-scale host risk ranking.

38%
NDCG@5 relative improvement

HeaRank’s gain over LightGBM Ranker for the highest-priority nodes.

64%
Failures captured in top risk tier

Share of future failures found within the top 5% of ranked nodes online.

0.881
30-day horizon AUC

Ranking performance under the longer labeling horizon.

What the paper found

Researchers from the Chinese Academy of Sciences, Tsinghua University, and StepFun argue that GPU reliability should be prioritized rather than precisely predicted. In production clusters, workload migration, non-stationary telemetry, low signal-to-noise ratios, and overlapping pre-failure and normal distributions make exact failure timing unreliable; even XGBoost, CNN, LSTM, Transformer, and Mixture-of-Experts models perform poorly, with the best top-K F1 reaching only 0.4837. The study focuses on Double Bit Errors and GPU Lost events, which account for 43.3% of service-impacting hardware incidents, and relates the problem to large-scale AI operations such as Meta’s LLaMA-3 training on 16,384 GPUs. The proposed HeaRank system reformulates reliability assessment as host-level risk ranking. Its compact MLP uses cumulative failure counts, recurrence and recency statistics, failure types, and hardware metadata instead of volatile time-series telemetry. On a production-scale, primarily NVIDIA cluster, HeaRank achieves an AUC of 0.834 and a 38% relative improvement in NDCG@5 over LightGBM Ranker. In six months of online deployment, it captured 64% of future failures within the top 5% of ranked nodes, compared with 21% for the incumbent health-score system. A 30-day labeling horizon raises AUC to 0.881, but the authors select a 7-day horizon because it better balances denoising with actionable maintenance scheduling. The central operational insight is to steer critical jobs toward lower-risk GPUs and reserve fragile nodes for tolerant workloads.

Original abstract

The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through an in-depth analysis of telemetry data from a productioncluster, we find that major GPU failures, including Double Bit Er-rors (DBEs) and GPU Lost events, exhibit strong stochasticity andlow signal-to-noise ratios in time-series telemetry, which makesconventional time-based prediction ineffective. This insight motivates a paradigm shift: instead of attempting topredict the absolute timing of a failure, we propose a more robustapproach focused on ranking nodes by their relative failure risk. Wepropose HeaRank (Health Rank), a Learning-to-Rank (LTR) frame-work that leverages stable historical failure patterns to computea global risk ranking of GPU nodes. Evaluated on a production-scale cluster with thousands of GPUs, HeaRank achieves an AUCof 0.83, significantly outperforming both heuristic baselines andstate-of-the-art ranking algorithms. In online deployment, HeaRanksuccessfully captures 64% of future failures within the top 5% ofranked nodes, compared to only 21% by the incumbent productionsystem. These results suggest that relative risk ranking can serveas a robust alternative in environments where absolute failure pre-diction is inherently limited. Our work highlights the importanceof risk-aware scheduling and proactive resource management inmodern GPU clusters.

Read the original paper

More in AI Hardware

Browse all 34 papers →