当前位置:首页 > 报告详情

High-Performance AI Fabrics Built on Open Standards_ Scaling 800G and 1.6T Infrastructures.pdf

上传人: S** 编号:1240930 2026-05-16 18页 1.27MB

1、High-Performance AI Fabrics Built on Open Standards:Scaling 800G and 1.6T InfrastructuresArt Fewell Technical Product Marketing,Hedgehog Inc.Arun Solleti Technical Product Marketing,Celestica Inc.AI CLUSTERS2Agenda01Challenges in Scaling High Performance AI Fabrics02Deployment needs and unified AI F

2、abric architecture03Optimizations for AI Training and Inference Fabrics 04AI Fabric Control,Contributions&Demo05Open Cluster Design for AI WG&How to participate3xPU#:1000s MillionsRack Power:kW MWSite Power:Grow to GWIndividual Cluster GrowthAI Fabrics Growth DemandsThroughput vs LatencyInterconnect

3、 vs ExternalPrecision vs EfficiencyTraining vs InferenceN America w/75%global AI super compute powerAPAC is fastest growing regionGrowth Across RegionsPower&Density+Complexity+Market GrowthAI Infra CAGR:30.4%until 2030.AI Server CAGR:39%until 2030.Global Market CapacityScaling AI Fabrics to Meet Glo

4、bal DemandRequirements for AI Fabric Deployments4Deploy high-performance,predictable AI clusters at scalecombining rapid,simplified setup with the flexibility to grow ScalabilityAI Training RA Max performance,massive scale,fault tolerance.AI Inference RA Low latency,multi-tenancy,operational simplic

5、ityPrescriptive,open-source,and OCP compliant design templates to accelerate AI cluster deployments CustomizeStandardizeOCP Open Cluster Design for AI5Individual Building NodesxPU-NodeStorageScale UpRackBuilding Block for Open Pod GroupOpen Pod Group of M xPUs(OPG-M)Multiple xPU-NodesStorage FabricS

6、cale Out Leaf SwitchOCP Open Rack v3 SpecBuilding Block for Open ClusterxPU Open Cluster of N-xPU(XOC-N)Scale Out SpineScale Out FabricScale Out CPUNetwork Services(Firewall)Building Block for AI ClustersOpen Cluster Network Fabric Unified Architecture6Control PlaneData PlaneAutomationPowered by Hed

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
1. **AI基础设施增长需求**:全球AI基础设施CAGR达30.4%(至2030年),AI服务器CAGR达39%,北美占全球75%算力,APAC增长最快。 2. **核心挑战**:需解决高密度、高功耗(kW→MW)、低延迟、可扩展性(千→百万xPU)及训练/推理差异化需求。 3. **开放架构设计**:基于OCP标准,采用Hedgehog AI Network(Kubernetes原生控制平面)与SONiC NOS,支持全生命周期自动化(Day 0-2)。 4. **优化方案**: - 训练:支持64-1024 xPU规模,CLOS/优化拓扑,双平面容错。 - 推理:低延迟ECMP、多租户VPC、网关服务(NAT/PAT)。 5. **开放协作**:Open Cluster Design for AI工作组推进标准化,2026年发布1.6T架构,2027年扩展xPU支持。
**AI Fabric如何扩展?** **低延迟如何实现?** **开源方案有哪些优势?**
客服
商务合作
小程序
服务号
折叠