Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
DGX agentarXiv:2504.11320v3 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires toke