1、Composable Memory Fabricsfor XPU-Scale AI Data ServicesGaurav Agarwal Senior Distinguished Engineer,MarvellVijay Ram Inavolu Senior Principal Engineer,MarvellAI CLUSTERSContextual background:AI serving workloadsMemory wallAI and memory wallAdvancement in compute has far outpaced memory(ref AI and me
2、mory wall)Memory constrained accelerators(XPU)struggle with sparsity,low arithmetic opsHigh costJohn HennesseyLLM serving costs up to 10X of traditional keyword search(John Hennessey)Under utilized XPUs,increasing complex pipelines(Agentic AI,MoE)Scaled datasets10M tokensOpenAIrefLarge sequences:100
3、K to 1M,even 10M tokensMassive embedding tables,highly dimensional vectors,e.g.3072(OpenAI)Huge models e.g.,GPT-4 1.8 trillion parameters 10X of GPT-3(ref)Efficiency needMcKinseyPower for AI is most capital intensive($1.3 trillion)after technology(ref McKinsey)Inferencing efficiency is crucial broad
4、 deployments Tiered Memory and StorageDeep cache and memory hierarchiesVaried NUMA distancesAccess methods-DMA,load/store,P2PAddressability:Byte/Block/Object/FileCompression,EncryptionPersistenceMulti-host sharingComputeCPUsGPUsDSAsFPGAsInterconnectsCoherent/non-coherentChannel and memory semanticsD
5、iverse workload-optimized topologiesHeterogenous SystemsArise because of multi-vendor staged HW procurement,domain-specific architecture trend,and rapid evolution of compute,memory,and fabricsPCIe/CXL AcceleratorStorageClassMemoryStoragewith computeMemory applianceComputational memoryHybrid switchin
6、gUALinkxPUxPUxPUDisaggregation adds considerable complexityData movement overheadsEnergy tax:up to 90%Bandwidth required:10 x 60 xPointers in near vs far memoryAccess latencies:more than 2xResponse time degradation:20 xCaching and prefetching limitationsEffectiveness:as low as 10%Reuse distance:over