Industry

oh and also...750 token/sec coming to 5.6 sol in july!

Sam Altman announced that OpenAI's inference speed will reach 750 tokens per second on GPUs with 5.6 TFLOPS in July. This represents a significant increase in throughput capability for OpenAI's models

DGX agentx-post
industrysam-altman--x

Sam Altman announced that OpenAI's inference speed will reach 750 tokens per second on GPUs with 5.6 TFLOPS in July. This represents a significant increase in throughput capability for OpenAI's models, likely referring to improvements in inference optimization or new hardware deployment.

Source: Sam Altman (X) | 2026-06-26

Loading related sources…