当前位置:首页 > 报告详情

HCiM:用于深度学习工作负载的无 ADC 混合模拟数字计算内存加速器.pdf

上传人: 芦苇 编号:651802 2025-05-01 23页 1.41MB

1、HCiM:ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning WorkloadsShubham Negi,Utkarsh Saxena,Deepika Sharma and Kaushik Roy22ndJanuary 2025 Introduction and Background Challenges Proposed Hardware Algorithm Co-design Approach Two Stage Quantization Hybrid Analog-Digital C

2、iM Accelerator Results Conclusion21BackgroundChallengesHCiMResultsConclusionIntroduction2Bring compute to the edgeSo,what is stopping us?Normalized Energy/MACDNN DataflowsSource:Sze V et al,Proceedings of IEEE,2017Source:IBMBreakdown of operation type across ML workloadsExploding computational compl

3、exity of Deep Learning modelsSource:OurWorldInData.org/artificial-intelligence Source:Magnet,ICCAD 2019Potential Solution:Compute in MemoryChallengesHCiMResultsConclusionIntroductionBackground34WLDBitwise NAND/NOR OutputBitwise NAND/NOR OutputDigital Compute In MemoryDigital Compute In Memory Digita

4、l CiM Bitwise multiplication operation followed by an accumulation in the peripherals Vector Matrix multiplication followed by reduction in adder tree+WLDBitwise NOR operationBitwise NOR operation+Shift/addWLDMUXMUXADCADCBitwise MVM OutputBitwise MVM OutputAnalog Compute In MemoryAnalog Compute In M

5、emoryAdder TreeAdder Tree123Chih et al.ISSCC 2021 Analog CiM Reduction performed in analog domain Multiple wordlines are turned on to perform bitwise MVM operation in Analog domainIntroductionBackgroundChallengesHCiMNAXFuture Work45Source:Ankit A et al,ASPLOS,2019PUMA:CiM Accelerator1234A(4 bit)5678

6、912345678912B(4 bit)xBit StreamIj=ViGijBit SlicesB1:0B3:201 10 11 1001 01 10 1100 01 10 1100 01 01 1001 01 01 1010 00 00 0001 01 01 0110 10 00 005 56 67 78 89 91 12 23 34 45 56 67 78 89 91 12 2MVMUMVMU0110A01000A11000A20000A30000A31000A21000A10110A012341234I1I2I3I4I1I2I3I4Divide the B matrix into bi

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文提出了一种名为HCiM的混合模拟-数字计算存储加速器,旨在为深度学习工作负载提供ADC-less(无需模数转换器)的解决方案。文章首先概述了深度学习模型中计算复杂性爆炸的问题,以及现有技术中ADC的能源和面积瓶颈。接着,文章介绍了一种两阶段量化方法,通过在训练过程中引入部分和量化,可以在不使用ADC的情况下实现CiM(计算存储)加速器。作者还提出了一种新的混合模拟-数字CiM宏,可以有效地处理量化后的部分和。实验结果显示,与传统的1位ADC相比,HCiM可以实现更高的准确度,同时降低能量延迟面积产品(EDAP)3.8倍;与BitSplitNet相比,HCiM在准确度上提高了2.5%,EDAP降低了3.8倍。在128x128的矩阵上,HCiM比使用7位ADC的基线能耗低28倍,比使用4位ADC的基线能耗低12倍。这些成果表明,HCiM在能量效率和准确性方面都有显著优势。
"混合模拟数字计算内存加速器如何工作?" "如何通过两阶段量化消除模拟CiM加速器中的ADC需求?" "HCiM加速器与现有技术相比有哪些优势和劣势?"
客服
商务合作
小程序
服务号
折叠