On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
arXiv:2602.01997v2 Announce Type: replace-cross Abstract: Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classific