当前位置:首页 > 报告详情

在资源受限的边缘计算平台上高效部署大型语言模型.pdf

上传人: 芦苇 编号:651849 2025-05-01 44页 12.78MB

1、2/16/251Efficient Deployment of Large Language Models on Resource Constrained Edge Computing PlatformsYiyu Shi,Ph.D.Professor,Dept.of Computer Science and Engineering,Site Director,NSF I/UCRC on Alternative and Sustainable Intelligent Computing,University of Notre Dame yshi4nd.edu11The Success of La

2、rge Language ModelsChemistryMedicineMathBusinessAnalyticsHosted on Cluster2“As models scale,they approach or surpass task-specific baselines,showing promise as universal systems for natural language understanding”-By Scaling Law from OpenAI22/16/252LLM is powerful,butOfflineData PrivacyAI Centraliza

3、tion(Fairness)CustomizationVision:LLM hosted on cluster can achieve many tasks,but is compromised by certain concerns:Offline Internet is unavailable/unstable,but real-time reaction is required(suicide detection,auto-drive)Data Privacy Medical history,personal informationAI Centralization Only large

4、 corps can own models,data,and computational resources(clusters)Customization LLM needs to adapt users with distinct situations33Edge-based LLM can be a solution“Data in local”“Model weights in local”“Free from Internet”“Customize the LLM via local data”LLM deployed on the edge device can avoid thes

5、e concerns.Microsofts Phi model,has successfully demonstrated the power of edge-friendly LLM442/16/253Gap Between LLM and Edge DevicesGAPLLM is growing much faster than the upgrade of edge devicesChallenges:Computation complexityMemory capacityEnergy efficiency55A Successful Edge LLM should be able

6、to Tradeoff:Use resource wisely among model weights and user data during training/inferencePersonalization:Generate user-preferred/related responseRobustness:Continuously growing performance over experienceHandle out-of-distribution scenarios 662/16/254Build Up Efficient LLM on Edge Devices Edge LLM

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要研究了在资源受限的边缘计算平台上高效部署大型语言模型(LLM)的方法。主要内容包括: 1. 模型设计:通过全面评估学习、模型权重和用户数据之间的权衡,为边缘LLM部署提供指导。 2. 数据选择:提出了一种基于质量指标的数据选择框架,以在边缘设备上维护高质量和紧凑的用户生成数据块。 3. RAG-CiM:通过非易失性计算内存在内存(NVCiM)加速检索增强生成(RAG),以优化LLM个性化。 4. NVCiM-PT:优化提示调整,一种基于NVCiM架构的LLM训练方法。 5. Tiny-Align:一种资源高效的跨模态(音频,文本)对齐方法,以实现LLM与用户的音频交互。 研究结果表明,适当的模型选择、数据选择和优化方法可以提高边缘LLM的性能,并使其更适合资源受限的边缘环境。
如何在边缘设备上高效部署大型语言模型? 边缘设备上的数据选择如何优化大型语言模型? 边缘计算如何实现大型语言模型的个性化?
客服
商务合作
小程序
服务号
折叠