Hardware
LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench (PKB) measures how they fail by benchma…
LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench (PKB) measures how they fail by benchmarking against 87 problems pulled from real codebases includi
LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench (PKB) measures how they fail by benchmarking against 87 problems pulled from real codebases including Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, NeMo-RL
Source: Together AI (X) | 2026-06-23