Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
arXiv:2603.17019v2 Announce Type: replace Abstract: A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only i