AMiD: Knowledge Distillation for LLMs with alpha-mixture Assistant Distribution
arXiv:2510.15982v3 Announce Type: replace-cross Abstract: Autoregressive large language models (LLMs) have achieved remarkable improvement across many tasks but incur high computational and memory cos