
Staff Engineer, Inference Optimizations
DigitalOcean seeks a Staff Engineer (Inference Optimizations) to lead benchmarking and performance work at the inference engine and GPU-kernel level. Responsibilities include attention-layer and kernel optimizations, memory/precision management, multi-node GPU parallelization, and deploying quantization techniques. Requires 5+ years in HPC or AI infrastructure and deep experience with NVIDIA/AMD GPUs, CUDA/ROCm, Triton/CUDA kernels. Compensation range $191,200–$239,000/year.