Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT
arXiv:2606.26861v1 Announce Type: new Abstract: Deploying large language models (LLMs) on Industrial Internet of Things (IIoT) edge devices demands extreme compression, yet existing structured pruning