In a distributed training environment, which NCCL primitive is required to combine gradient updates from all GPUs into a single synchronized result across the cluster?
Community Answers
Sign in to open profiles and full community answers.
No community answers yet. Be the first to submit one.