Optimizing LLM Inference: Disaggregated Serving, PD Protocol, & KV Pinning

Video by PyTorch via YouTube
Optimizing LLM Inference: Disaggregated Serving, PD Protocol, & KV Pinning

At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at Mistral AI and vLLM maintainer, will present joint work with Amazon Web Services (AWS) and Red Hat on how disaggregated serving in vLLM has evolved to support the latest generation of hybrid models.

Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60

Source