Hardware
New in Together GPU Clusters: Reliability and control for production GPU clusters
Together GPU Clusters now includes a suite of resilience and operational‑control features designed to address common large‑scale failure modes and team‑scaling challenges. The platform‑health improvem
Together GPU Clusters now includes a suite of resilience and operational‑control features designed to address common large‑scale failure modes and team‑scaling challenges. The platform‑health improvements—passive health checks that continuously monitor real workloads, auto‑node repair, and the Slinky 1.0 runtime—detect and recover from GPU‑bus failures, Xid errors, thermal throttling, and scheduler leaks, while the operational‑control upgrades add a detailed cluster view, external OIDC authentication, customizable startup scripts, and an acceptance‑test opt‑out to give teams greater visibility and flexibility. Together, these changes enable faster failure detection, cleaner recovery, and a more manageable workflow as organizations grow beyond a single‑admin‑kubeconfig model.
Related
- Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams
- GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures
- Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
- Don't Let a Few Network Failures Slow the Entire AllReduce
Source: Together AI Blog | 2026-07-15