Video by PyTorch via YouTube

Elastic Expert Parallelism in vLLM lets you add or remove GPUs from an active Mixture-of-Experts deployment during traffic with minimal interruption to serving and minimal downtime.
At #PyTorchCon North America 2026, Itay Alroy of NVIDIA will present “Elastic Expert Parallelism in vLLM,” covering the architecture, key implementation details, open challenges, and future roadmap for Elastic EP.
Alroy will also examine what happens when EP size changes and how NIXL EP enables grow/shrink under live traffic.
Register for PyTorch Conference North America 2026: https://hubs.la/Q04v4SL60