Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
arXiv:2606.21428v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small su