POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
DGX agentarXiv:2604.16583v1 Announce Type: new Abstract: Edge deployment of large language models (LLMs) increasingly relies on libraries of lightweight LoRA adapters, yet GPU/DRAM can keep only a small reside