Model Releases
GPT-5.6 found optimizations that 'reduced end-to-end serving costs by 20%' for OpenAI to serve that model Presumably that's billions of doll…
GPT-5.6 found optimizations that 'reduced end-to-end serving costs by 20%' for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Codex analysed product
GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation …
Source: Simon Willison (X) | 2026-07-30