Efficient Pre-Training of LLMs through Truncated SVD Layers
arXiv:2605.28573v1 Announce Type: cross Abstract: The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal