WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
DGX agentarXiv:2607.02391v1 Announce Type: cross Abstract: Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requi