Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment
arXiv:2607.08268v1 Announce Type: new Abstract: High-volume structured extraction pays a large model's latency on every item, so distilling the task into a small on-device model is attractive: compara