当前位置:首页 > 报告详情

Neuron 的性能工程:如何使用 NKI 优化您的 LLM.pdf

上传人: 明**** 编号:1013569 2025-12-21 17页 542.88KB

1、 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.A I M 4 1 4Performance engineering on Neuron:How to optimize your LLM with NKIScott PerryPrincipal Solutions Architect,AI/ML PerformanceAnnapurna Labs,AWSSadaf Rasoo

2、lSolutions Architect,AI/ML PerformanceAnnapurna Labs,AWS 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.Innovating at the silicon levelInnovating at the silicon level3AWS TrainiumAWS InferentiaAWS AI Chips 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.AWS AI

3、Chipsfor Generative AIAWS Inferentia AWS Inferentia2 AWS Trainium AWS Trainium2Deep learning modelsMedium to large-scale inferenceLLMs,multi-modal modelsMedium to large-scale training and inference:LLMs,multi-modal modelsTraining and inference for Gen AI modelsAWS Trainium3AWS AI ChipsNext-gen agent

4、ic,reasoning,and video generation applications 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.NeuronCore ArchitectureHBMNeuronCoreGPSIMD EngineScalar EngineVector EngineTensor EngineDMA EnginesPSUMSBUFHost(CPU)Memory 2025,Amazon Web Services,Inc.or its affiliates.All rights reser

5、ved.Memory HierarchySRAMAccelerator HBMHost Memory-Size:MBs-Bandwidth:10TB/s-Size:10s GBs-Bandwidth:TB/s-Size:10s GBs TBs-Bandwidth:GB/s 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.How do we improve performance?Pipeline operations Minimize data movement Maximize data throughpu

6、t Collectives time compute/data ops timecompute boundPerformanceArithmetic Intensity(ops/byte)2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.ML DevelopersData ScientistsPerformance EngineersNeuron Developer Stack 2025,Amazon Web Services,Inc.or its affiliates.All rights reserved.

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
根据报告的内容,全文主要内容概括如下: - **AWS Trainium和Inferentia芯片**:介绍AWS Trainium和Inferentia芯片,用于深度学习模型的中到大规模训练和推理,以及生成式AI应用。 - **NeuronCore架构**:详细描述NeuronCore的架构,包括HBM、GPSIMD、Scalar、Vector、Tensor引擎和DMA引擎等。 - **性能优化**:强调通过优化数据移动、提高数据吞吐量和减少集体操作时间来提升性能。 - **Neuron Developer Stack**:介绍Neuron Developer Stack,包括NKI(Neuron Kernel Interface)和Neuron Kernel Library,用于编写和优化内核。 - **NKI库**:发布NKI库,提供预优化的内核,支持密集模型、混合专家模型等,并开源在GitHub上。 - **性能提升案例**:通过Qwen3模型案例展示使用NKI内核前后的性能提升,吞吐量提高2.3倍,延迟降低。 - **会议日程**:列出一系列与AWS Trainium和Inferentia相关的会议和研讨会,包括性能工程、模型优化、成本效益等主题。
"NKI优化LLM,性能翻倍?" "AWS Trainium,如何打造高性能模型?" "AI芯片新篇章,NKI库揭秘!"
客服
商务合作
小程序
服务号
折叠