Protocol

NCCL

NVIDIA Collective Communications Library

Lab ready Full lab

NVIDIA Collective Communications Library implementing allreduce, broadcast, and related GPU collectives over NVLink, PCIe, and RDMA. Topology detection and ring/tree algorithms dominate training performance at scale. It is the default collective layer for many PyTorch/TensorFlow multi-GPU jobs.

Domains

Tags

collectivesgpuallreducetraining

Sources

  • NVIDIA NCCL documentation

Related protocols