
Senior HPC AI Cluster Engineer
Role to design, implement, tune and maintain large-scale HPC/AI clusters (monitoring, logging, alerting), manage job scheduling/orchestration (e.g., Slurm, K8s), develop CI/CD and automation tooling, and troubleshoot from bare metal to application. Requires 8+ years experience and knowledge of CPUs/GPUs, high-speed interconnects, Linux/Windows internals, storage solutions (Lustre, GPFS, Weka.io), networking protocols and virtualization/cloud platforms.
