
Senior Site Reliability Engineer, DGX Cloud
Join NVIDIA's DGX Cloud team to build, implement and support operational and reliability aspects of large-scale Kubernetes clusters focused on performance at scale. The role requires defining SLOs/SLIs, operating GPU workloads across AWS/GCP/Azure/OCI and private clouds, building observability stacks, automating infrastructure, participating in on-call rotation, and leading incident triage; 10+ years of production operations experience and expertise in Kubernetes, Linux, cloud platforms, and SRE practices are required.
