1、GPU Communication Libraryin Meta-Scale AI ClustersJames Hongyi ZengMetaAI CLUSTERSMetas AI Clusters Getting Big 1GW,Announced 2026Lebanon,IndianaUp to 5GW,ETA 2028HyperionRichland Parish,Louisiana1GWPrometheusNew Albany,OhioDiverse Communication NeedsLLM StagesExample ChallengesCritical CommsExample
2、 OSS FrameworksExample OSS Comms LibrariesPre-TrainingMaximize massive throughput across thousands of GPUsFSDP(AllGather/ReduceScatter)torchtitan,Megatron-LMNCCL,RCCLPost-Training(RLHF)Rapid weight shipping from Learners to ActorsP2P(Send/Recv)torchforge,verlRay RPCInferenceKV Cache shipping for Pre
3、fill-Decode disaggregationP2P(Send/Recv)Expert Parallelism(AlltoAll)vLLM,SGLangDeepEP,Mooncake TEPre-training:Fault Tolerance Frequent failures are inherent risk for large scale synchronous pre-training18 min at 100K GPUs10 min restart time8 min effective training!Diverse set of interruptionUnavoida
4、ble Hardware FailuresPre-training:Fault Tolerance Solution:Parallelism Aware Fault ToleranceDivide GPUs into multiple replicasSame rank in each replica communicates with AllReduceDynamically scale up and downMeta Collective Communication Library(MCCL)Fast recoveryOne replica down,the rest continuesF
5、ast fault localizationSpare machines to restart replicaInference:Device Centric Communication High performance Mixed-of-Experts(MoE)is important for InferenceAll2All communicationDevice Centric CommunicationHost-based GPUDirect Async(aka IBRC)Device-native GPUDirect Async RDMA(aka IBGDA)Existing Lib
6、rariesNVSHMEMNCCL GPU Initiated Networking(GIN)Host-based GPUDirect Async(aka IBRC)Device-native GPUDirect Async RDMA(aka IBGDA)Inference:Device Centric Communication Pipes Work in progress,Open sourcing soon Device-native communication framework for writing c