
Senior Engineer, Inference Data Plane
Hands-on technical leader responsible for architecting and delivering distributed inference hosting on Kubernetes (using frameworks like llm-d, vLLM, Ray Serve), optimizing performance (KV-cache, batching, routing), defining SLOs, operating high-scale services, and contributing to upstream open-source inference projects. Requires strong software engineering (Go/Python, gRPC), distributed systems expertise, and experience with LLM inference engines and optimizations.








