1、High-Performance AI Fabrics Built on Open Standards:Scaling 800G and 1.6T InfrastructuresArt Fewell Technical Product Marketing,Hedgehog Inc.Arun Solleti Technical Product Marketing,Celestica Inc.AI CLUSTERS2Agenda01Challenges in Scaling High Performance AI Fabrics02Deployment needs and unified AI F
2、abric architecture03Optimizations for AI Training and Inference Fabrics 04AI Fabric Control,Contributions&Demo05Open Cluster Design for AI WG&How to participate3xPU#:1000s MillionsRack Power:kW MWSite Power:Grow to GWIndividual Cluster GrowthAI Fabrics Growth DemandsThroughput vs LatencyInterconnect
3、 vs ExternalPrecision vs EfficiencyTraining vs InferenceN America w/75%global AI super compute powerAPAC is fastest growing regionGrowth Across RegionsPower&Density+Complexity+Market GrowthAI Infra CAGR:30.4%until 2030.AI Server CAGR:39%until 2030.Global Market CapacityScaling AI Fabrics to Meet Glo
4、bal DemandRequirements for AI Fabric Deployments4Deploy high-performance,predictable AI clusters at scalecombining rapid,simplified setup with the flexibility to grow ScalabilityAI Training RA Max performance,massive scale,fault tolerance.AI Inference RA Low latency,multi-tenancy,operational simplic
5、ityPrescriptive,open-source,and OCP compliant design templates to accelerate AI cluster deployments CustomizeStandardizeOCP Open Cluster Design for AI5Individual Building NodesxPU-NodeStorageScale UpRackBuilding Block for Open Pod GroupOpen Pod Group of M xPUs(OPG-M)Multiple xPU-NodesStorage FabricS
6、cale Out Leaf SwitchOCP Open Rack v3 SpecBuilding Block for Open ClusterxPU Open Cluster of N-xPU(XOC-N)Scale Out SpineScale Out FabricScale Out CPUNetwork Services(Firewall)Building Block for AI ClustersOpen Cluster Network Fabric Unified Architecture6Control PlaneData PlaneAutomationPowered by Hed