
Senior HPC AI Cluster Engineer
Role involves architecting, deploying, and tuning large-scale HPC and AI clusters (CPU/GPU), managing job scheduling and orchestration (e.g., Slurm, Kubernetes), developing CI/CD pipelines and automation, deploying monitoring/logging/alerting, troubleshooting from bare metal to application level, and supporting R&D and POCs. Requires strong Linux/Windows, networking, storage, scripting, and virtualization knowledge.












