
Senior HPC Cluster Engineer - AI, ML
NVIDIA is hiring a Senior AI/ML HPC Cluster Engineer to lead day-to-day operations, reliability, and performance optimization of on-premises and cloud GPU clusters. The role requires a minimum of 5 years' experience with large-scale compute infrastructure and hands-on skills with schedulers (e.g., Slurm, K8s, PBS), Linux administration, configuration management, container technologies, Python/bash scripting, and MPI-based workflows.











