
Senior System Software Engineer, Agentic Inference - Dynamo
Work on Dynamo's open-source GPU-accelerated inference stack, extending disaggregated serving across vLLM, SGLang, and TensorRT-LLM to support agentic inference (long-horizon reasoning, tool calling, stateful multi-turn execution). Responsibilities include inference-state management (KV/prefix caches, memory/storage hierarchies), building distributed frontends, optimizing throughput and latency, and contributing to upstream API compatibility. Requires 10+ years' experience, strong Rust and Python skills, and expertise in LLM inference and GPU optimization.