Video by PyTorch via YouTube

Communication overhead remains one of the largest bottlenecks in scaling distributed training.
In an upcoming session at PyTorch Conference North America, Rishi Sinha from AMD shares how integrating work from the RCCL team enables GPU-initiated networking (GIN) directly within TorchTitan, PyTorch’s distributed training framework.
By allowing the GPU kernel itself to issue network puts to remote GPUs, this approach improves all-to-all kernel performance by up to 30 percent and delivers a 12 to 15 percent boost in end-to-end performance.
Join us in San Jose on October 20-21 to explore the future of distributed training infrastructure: https://hubs.la/Q04v4SL60