
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Role owns full lifecycle of GPU compute clusters (procurement, provisioning, config management, monitoring, deprecation) across heterogeneous Linux environments. Requires 5+ years administering large-scale HPC/ML clusters, strong Linux and scripting skills, experience with Slurm, Ansible, storage solutions (NFS/Lustre/WekaFS), container runtimes, and high-speed networking technologies.