
NCX Senior Engineer
Hands-on senior engineering role working with NVIDIA Cloud Partners to build Day 2 operational capabilities for large-scale GPU-accelerated clusters. Responsibilities include observability, continuous infrastructure validation, automated detection and remediation, lifecycle and fleet administration (drivers/firmware, OS patching, Kubernetes node maintenance), and translating NVIDIA reference architectures into repeatable production operating practices. Requires 8+ years experience in infrastructure/SRE/DevOps, strong Linux and Kubernetes expertise, production observability, automation skills (Python/Go/shell), and networking/storage troubleshooting.









