Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod
Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence.
Key points
- Amazon SageMaker HyperPod is a purpose-built AI service that simplifies the management of large-scale compute clusters for gen AI workloads.
- In this post, we present a reference architecture for building a multi-tenant environment on Amazon SageMaker HyperPod with EKS.
- This architecture uses AWS IAM Identity Center for centralized authentication, per-team SageMaker AI domains for a tailored user experience, Kubernetes namespaces for workload isolation, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility and chargeback.
- Figure 1: High-level multi-tenant architecture for two teams sharing one HyperPod EKS cluster
Sources (1)
- [1]Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPodAWS Machine Learning Blog · Oct 8, 04:20 PM
Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence.
Amazon SageMaker HyperPod is a purpose-built AI service that simplifies the management of large-scale compute clusters for gen AI workloads.
Extractive summary: sentences quoted from the sources.