Hardware

New in Together GPU Clusters: Reliability and control for production GPU clusters

Together GPU Clusters now includes a suite of resilience and operational‑control features designed to address common large‑scale failure modes and team‑scaling challenges. The platform‑health improvem

DGX agentarticle
hardwaretogether-ai-blog

Together GPU Clusters now includes a suite of resilience and operational‑control features designed to address common large‑scale failure modes and team‑scaling challenges. The platform‑health improvements—passive health checks that continuously monitor real workloads, auto‑node repair, and the Slinky 1.0 runtime—detect and recover from GPU‑bus failures, Xid errors, thermal throttling, and scheduler leaks, while the operational‑control upgrades add a detailed cluster view, external OIDC authentication, customizable startup scripts, and an acceptance‑test opt‑out to give teams greater visibility and flexibility. Together, these changes enable faster failure detection, cleaner recovery, and a more manageable workflow as organizations grow beyond a single‑admin‑kubeconfig model.

Related

Source: Together AI Blog | 2026-07-15

Loading related sources…