
Senior AI Performance and Efficiency Engineer
Role responsible for improving ML training and inference efficiency across GPU clusters by debugging and optimizing software and infrastructure, building tools and frameworks to detect bottlenecks, collaborating with researchers across domains (LLMs, robotics, autonomous vehicles, video), and working with technologies such as CUDA, NCCL, Nsight, distributed training, cloud platforms, and fast distributed storage.








