当前位置:首页 > 报告详情

HotChips_Boqueria_Presentation_v15.pdf

上传人: 2*** 编号:136912 2023-08-03 19页 2.18MB

1、BoqueriaRobert Beachler VP of Product/Hardware EngineeringDr.Martin Snelgrove Co-founder and CTOCopyright 2022 UNTETHER AI Corp.A Brief History of the Current AI Summer201220162020201820222014Deepmindacquired by GoogleAlphaGo beats Lee SedolUntether AI founded in TorontorunAI200 introducedFirst at-m

2、emory inference accelerator500 INT8 TOPs200MB SRAM8 TOPs/WTSMC 16nmBoqueria introduced2 PetaFlops FP8238MB SRAM30 TFLOPs/WTSMC 7nmCopyright 2022 UNTETHER AI Corp.AI Inference Presents 3 Key Challenges to Chip MakersIncreasing computational and power requirementsScalability and flexibility for changi

3、ng NN landscapeAccuracy loss costs$millions and risks livesThe Computational Limits of Deep LearningNeil C.Thompson1,Kristjan Greenewald2,Keeheon Lee3,Gabriel F.MansonFirst-Generation Inference Accelerator Deployment at FacebookNHSTA report,June 2022 for July 2021 to May 2022Model CategoryModel Name

4、Model Size(Mparams)RecommendationLess complex70,000More complex100,000Computer VisionResNeXt101-32x4-4844RegNetY700FBNetV3 based model28.6Video UnderstandingResNeXt3D based58NLPXLRM-R558Copyright 2022 UNTETHER AI Corp.Architecting an AI Inference AcceleratorPower-efficient throughput is required to

5、meet NN compute demandData movement is the costliest part of inference 90%of energy consumptionData movement is different between training and inferenceOptimizing compute architecture to minimizing distance travelled results in inference-specific AI acceleratorsProper level of granularity to create

6、a scalable compute architectureRight balance between coarse-grained and fine-grained approachDont over-fit for a particular application/NNUtilize the most efficient datatype for a given application and accuracy requirementsA mixture of datatypes provides the best resultsDesigning Energy-Efficient Co

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要介绍了Boqueria公司的一款AI推理加速器,该加速器采用at-memory计算架构,以能量效率和大规模并行处理为特色。关键数据包括:Boqueria加速器提供2 PetaFLOPs的计算能力,30 TFLOPs/W的能效,以及1,458个RISC-V核心。它支持多种数据类型,如FP8和BF16,以实现精确度和能效的平衡。该加速器还具备灵活的计算架构,可在不同的神经网络架构下扩展和调整。此外,Boqueria加速器通过优化数据移动和处理,实现了高效的能量消耗和计算吞吐量。与传统的GPU相比,Boqueria加速器在性能和能效方面具有显著优势,例如,在某些模型上,其性能提高了5倍,能效提高了7倍。
"Boqueria如何实现高效的AI推理加速?" "FP8数据类型在AI推理中的优势是什么?" "imAIgine SDK如何优化AI模型推理性能?"
客服
商务合作
小程序
服务号
折叠