Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
DGX agentarXiv:2510.03243v3 Announce Type: replace-cross Abstract: Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that