Collective communicationPurlin targets GPU collectives by splitting orchestration from the datapathA Stanford University and NVIDIA arXiv paper attacks a real distributed inference bottleneck, not the usual benchmark confetti cannon.NyxOct 1, 20264 min readRead the story