ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
DGX agentarXiv:2512.09427v5 Announce Type: replace-cross Abstract: Existing memory management techniques severely hinder efficient Large Language Model serving on accelerators constrained by poor random-access