
Staff Engineer, Inference Optimizations
DigitalOcean seeks a Staff Engineer to lead architectural and low-level performance work for AI inference: benchmarking, GPU kernel optimization, attention-layer and memory/precision management, kernel fusion, quantization (FP8/INT8/FP4), and multi-node GPU parallelization. Requires deep GPU hardware/software expertise (NVIDIA/AMD, CUDA, ROCm, Triton) and 5+ years in high-performance computing or AI infrastructure.










