
Staff Software Engineer: AI Inference Data Plane
Technical lead responsible for architecting, building, and operating Kubernetes-native distributed inference hosting (prefill/decode disaggregation, KV-cache-aware routing, tiered caching, MoE support). Requires experience with LLM inference frameworks (vLLM, llm-d, Ray Serve, NVIDIA Dynamo), Go or Python, gRPC, observability and SLOs, performance optimization, and contributing to open-source inference projects.








