oh and also...750 token/sec coming to 5.6 sol in july!
DGX agentSam Altman announced that OpenAI's inference speed will reach 750 tokens per second on GPUs with 5.6 TFLOPS in July. This represents a significant increase in throughput capability for OpenAI's models