TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
arXiv:2605.17821v1 Announce Type: cross Abstract: Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster