Hardware
Understanding GPU Inference Workloads [D]
Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services l
Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services like runpod or vast.ai, your perspective is extremely valuable. Please share your experience in the comments here or by DMing me. I've also made a 2 minute survey form that I would really appreciate if you could fill out. DM me for the link. Thank you! submitted by /u/chinmaydagod [link] [comments]
Related
- I am absolute clueless about online GPU rent and setup image generation, need some advice from seniors.
- GPU Compass – open-source, real-time GPU pricing across 20+ clouds [P]
- Rewriting model inference with CUDA kernels: the bottleneck was not just GEMM [P]
- Anyone using Tensordock GPU instances and having problems with failing VM’s [D]
- An open handbook on LLM inference at scale (GPU internals, KV cache, batching, vLLM/SGLang/TensorRT-LLM) [P]
Source: r/MachineLearning | 2026-07-26