Applications
The production platform for open-weight AI inference
OpenAI has updated its inference platform to give users full control over performance, cost, and quality without building their own stack—models go live in minutes and support multiple deployments beh
OpenAI has updated its inference platform to give users full control over performance, cost, and quality without building their own stack—models go live in minutes and support multiple deployments behind a single stable endpoint with canary, blue‑green, rolling updates, auto‑rollbacks, A/B and shadow testing, and autoscaling across regions. The platform also offers a closed‑beta custom training service for full‑weight and LoRA reinforcement learning and supervised fine‑tuning, producing checkpoints that can be deployed directly to production. Open‑weight models now match closed‑model quality at a fraction of the cost, providing teams with complete control over functionality, performance, and proprietary data integration.
Related
- Our cofounder @the_bunny_chen joined @GregorVand to talk about open models, production inference, and how RFT is unlocking model customizati…
- As open models get stronger, more workloads move into the competitive inference market. That pushes the real fight toward speed, cost, relia…
- Move from test to production by running high-performance inference directly on Foundry. At #MSBuild, we demoed an end-to-end workflow showin…
Source: Together AI Blog | 2026-07-23