Video by PyTorch via YouTube

SGLang-Diffusion is a high-performance serving framework for diffusion models, designed for both large-scale offline generation and latency-sensitive real-time inference. As diffusion workloads expand from image and video generation to diffusion-based world models, serving systems must handle very different constraints: maximizing GPU throughput for batch/offline jobs, while maintaining tight end-to-end latency for interactive and closed-loop environments.
At PyTorch Conference North America 2026, Yihao Wang and Kevin Mi of RadixArk will present the design of SGLang-Diffusion and the PyTorch-based inference optimizations behind it, including efficient request scheduling, memory management, batching strategies, and execution-path optimizations for diffusion pipelines.
They will discuss how serving diffusion models differs from serving LLMs, why existing inference systems are often insufficient for iterative denoising workloads, and how a unified framework can support both offline generation and real-time world-model inference.
Join us in San Jose, October 20-21: https://hubs.la/Q04v4SL60
#PyTorchCon