
Senior Deep Learning Sofware Infrastructure Engineer
Role to design, harden, and scale deep learning infrastructure and libraries for large-scale GPU training (data loaders, distributed training, scheduling, monitoring). Requires 12+ years experience in high-performance distributed systems, strong Python skills, and deep knowledge of PyTorch, DDP/FSDP, NCCL, datacenter networking, parallel filesystems, and schedulers.











