
Senior Deep Learning Software Engineer, Inference and Model Optimization
Role entails training, developing, and deploying generative AI models using NVIDIA's AI stack; extracting model graphs from PyTorch, developing inference optimization techniques (sharding, efficient kernels, quantization, sparsity), and profiling GPU performance. Requires a Masters/PhD or equivalent experience, 8+ years in deep learning, strong software design and proficiency in Python, PyTorch, and GPU/kernel tooling.










