Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
DGX agentarXiv:2607.10183v2 Announce Type: replace-cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory ca