The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
DGX agentarXiv:2607.23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is th