Video by CNCF [Cloud Native Computing Foundation] via YouTube

Join Alexa Griffith of Red Hat at KubeCon + CloudNativeCon North America, November 9-12 in Salt Lake City, Utah.
In “Cache Me If You Can: LLM Inference with llm-d (with fun drawings),” Alexa will explore how to make LLM inference on Kubernetes more efficient and cost-aware.
Learn how llm-d uses prefix cache-aware routing, disaggregated prefill and decode, and flow control to improve performance as inference workloads scale. Expect clear explanations, honest tradeoffs, and yes, fun drawings.
Explore the session:
https://bit.ly/4y5Gmqa
Register:
https://bit.ly/3BDI5XL