L20-Edu-135M: An Auditable Single-GPU Study of Data-Efficient Small Language Modeling
DGX agentarXiv:2606.22189v1 Announce Type: new Abstract: Small language models are cheap to serve and feasible on local hardware, but strong public 135M-class systems are commonly trained with hundreds of bill