TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
AuthorsShujie Han, Feng Jiang, Patrick P. C. Lee, Xiao Zhang, Zhijie Huang, Nannan Zhao, Xiaonan Zhao, Lichen Pan
Resources
TierCheck makes LLM training more resilient by keeping quick-to-recover checkpoints close to the GPUs while asynchronously backing up full states to durable storage.
Key results
On a 16-GPU A800 testbed, TierCheck reduces checkpointing overhead versus CheckFreq, Gemini, and DataStates-LLM.
Average checkpointing time per iteration on the 20B-model experiments.
TierCheck accelerates recovery by using Tier-1 local memory and differential replay after software failures.
TierCheck speeds up recovery by restoring from Tier-2 peer memory after single-node failures.
Even for rack failures, TierCheck remains faster by falling back to Tier-3 remote persistent storage.
The system scales to models up to 40B parameters without accuracy degradation.
What the paper found
TierCheck targets a concrete bottleneck in large language model training: failures are not uniform, yet most checkpointing systems treat them as if they were. The paper shows that in production LLM runs, GPU and node crashes dominate, while rack outages are rare but catastrophic, and proposes a three-tier checkpoint hierarchy that matches that failure spectrum. Tier-1 keeps lightweight differential checkpoints in local volatile memory for software failures, Tier-2 asynchronously mirrors them to a physically adjacent peer node for single-node failures, and Tier-3 migrates base checkpoints to remote persistent storage for rack-level recovery. The novelty is not just placement, but lifecycle control: TierCheck intercepts DeepSpeed and PyTorch serialization to capture in-memory base checkpoints with zero-copy, compresses frequent differentials with adaptive INT8 quantization for small tensors and threshold-based Top-K-style sparsification for large tensors, then replays recovery using a fused multi-step operator that processes batches of differential checkpoints in one pass. A watermark-driven global reclamation protocol prevents premature deletion until a base checkpoint is durably committed cluster-wide. On a 16-GPU A800 testbed, TierCheck cuts checkpointing overhead by 62.8-82.7% versus CheckFreq, Gemini, and DataStates-LLM, reduces checkpointing time to 3.1-3.2 s, and speeds recovery by 3.6-6.9× for software failures, 1.8-3.5× for node failures, and 1.3-2.5× even for rack failures. It also preserves convergence on GPT2 124M trained on WikiText-2 and scales to 40B-parameter models without accuracy degradation.
Original abstract
Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster-wide outages. Existing checkpointing systems rely on monolithic, single-tier storage backend, forcing a trade-off between state-saving overhead and recovery speed. We propose TierCheck, a cluster-aware tiered checkpointing system that aligns storage placement with failure heterogeneity. TierCheck adopts a three-tier design that maintains lightweight differential checkpoints in local and peer memory for fast localized recovery, while asynchronously migrating heavyweight base checkpoints to remote persistent storage. It also ensures strict global consistency across tiers without stalling training, and achieves fast cluster-aware checkpoint restoration during recovery. Evaluations on models up to 40 billion parameters show that TierCheck achieves low training overhead, reduces end-to-end checkpointing time to under 10s, and supports high-frequency checkpointing, ultimately striking an optimal balance between low-overhead persistence and fast recovery.
Read the original paper