LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
arXiv:2605.01058v1 Announce Type: cross Abstract: Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; ye