YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition
arXiv:2606.05868v1 Announce Type: new Abstract: Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory