Best practices, reference architectures, and examples for distributed AI training and inference on AWS.
-
Updated
Sep 2, 2026 - Shell
Best practices, reference architectures, and examples for distributed AI training and inference on AWS.
This collection of helper scripts/ and guides for AWS SageMaker HyperPod and ParallelCluster makes it easy to get started with large-scale distributed training on Slurm-based HPC clusters and Kubernetes-based EKS clusters, as well as AI/ML model inference deployment.
Infrastructure deployment automation of SageMaker HyperPod clusters based on EKS and SLURM orchestration and Protein Language ESM-2 model training job definitions including NVIDIA BioNemo
Experimental scripts for Amazon SageMaker HyperPod
To associate your repository with the hyperpod topic, visit your repo's landing page and select "manage topics."