Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM
DGX agentTLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The performance is 2-3