
Staff Engineer, Inference Optimizations
Lead technical work on benchmarking and optimizing inference engines and GPU kernels to maximize throughput and minimize latency for large models. Responsibilities include attention-layer and memory optimizations, quantization (FP8/INT8/FP4), kernel fusion, multi-node GPU parallelization, and advising on GPU hardware/software stacks (CUDA, ROCm, Triton). Requires 5+ years in high-performance computing or AI infrastructure. Salary range provided.