This case study details the deployment of a cloud High-Performance Computing (HPC) cluster for artificial intelligence and machine learning research at a leading FAANG company. The objective was to scale the AI/ML research infrastructure to accommodate hundreds of researchers and thousands of GPUs, addressing the limitations of on-premises HPC clusters, which were costly and slow to expand. The solution involved the implementation of a customized HPC Slurm cluster on AWS, utilizing ParallelCluster as the foundational layer. This deployment included secure access, two-factor authentication, Unix user management, and multi-tenant support, integrating S3 pipelines and multiple FSx for Lustre file systems. An Azure HPC cluster was later added, creating a unified multi-cloud platform optimized for GPU-intensive workloads. The resulting environment supported over 500 researchers and managed extensive data, significantly enhancing collaboration and reducing experiment cycle times. The success of this deployment influenced AWS's subsequent enhancements to ParallelCluster.