Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
arXiv:2504.08791v3 Announce Type: replace-cross Abstract: On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low thr