Hardware
How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure
A combined architecture using KAI Scheduler and vCluster enables multiple teams to run fully isolated Kubernetes tenant clusters on a shared GPU node, each with its own control plane, RBAC, CRDs, and
A combined architecture using KAI Scheduler and vCluster enables multiple teams to run fully isolated Kubernetes tenant clusters on a shared GPU node, each with its own control plane, RBAC, CRDs, and cluster‑admin access.
KAI Scheduler provides topology‑aware hierarchical scheduling, per‑team quotas, dynamic allocation, and integrates with the NVIDIA GPU Operator, supporting large‑scale operation across thousands of nodes.
vCluster virtualizes separate Kubernetes clusters for each team, ensuring logical separation while sharing underlying hardware, thereby maximizing infrastructure efficiency without needing to split GPUs physically.
Related
- Running Large-Scale GPU Workloads on Kubernetes with Slurm
- Get Real-Time Visibility into GPU Usage Across Kubernetes Clusters
- Achieving Peak System and Workload Efficiency on NVIDIA GB200 NVL72 with Slurm Block Scheduling
Source: NVIDIA Developer | 2026-08-03