当前位置:首页 > 报告详情

探索基于 HLS 的自动化硬件生成的代码语言模型:基准、基础设施和分析.pdf

上传人: 芦苇 编号:651844 2025-05-01 33页 1.59MB

1、1Exploring Code Language Models for Automated HLS-based Hardware Generation:Benchmark,Infrastructure and Analysis ASP-DAC 2025Jiahao Gai,Hao(Mark)Chen,Zhican Wang,Hongyu Zhou,Wanru Zhao,Nicholas Lane,Hongxiang FanOutlineI.IntroductionII.DatasetIII.ModelIV.FrameworkV.EvaluationVI.Discussion23The Era

2、of Generative AILLM-assisted code generation:Github Copilot1,Deepminds AlphaCode2Over 50 pre-trained models and more than 170 programming language datasets releasedAutomated Hardware Design Generation:Verilog,SystemVerilog1 Chen,Mark,et al.Evaluating large language models trained on code.arXiv prepr

3、int arXiv:2107.03374(2021).2 Li,Yujia,et al.Competition-level code generation with alphacode.Science 378.6624(2022):1092-1097.Challenge 1:Data Availability of HDLC+=40.52 times HDLPython=2.26*104 times HDL4Challenge 2:Difficulty in Transferring Pretrained KnowledgeMost code LLMs pre-trained on softw

4、are programming languageDifferent from HDL5Challenge 3:Cost of GenerationHDL implementations require 34 times more tokens than HLS6A Code LLM for HLS GenerationChallenge 1&2:HLS shares main semantic/syntax with C/C+,which makes knowledge transfer possible and reduces dataset requirementsChallenge 3:

5、HLS generation is more cost-efficient at inference timeDataset+Model+Generation Framework7Research QuestionsWhether the existing public data is enough for the training HLS-Gen LLM?What performance can be achieved using existing public data?Can advanced techniques,such as CoT,help HLS-Gen?8OutlineI.I

6、ntroductionII.DatasetIII.ModelIV.FrameworkV.EvaluationVI.Discussion9Format of DatasetInput:Natural language description from developerOutput:HLS design10Dataset Collection52 designs,42000 HLS programs from HLSyn3 and ML4Accel4113 https:/ https:/ Thakur,Shailja,et al.Verigen:A large language model fo

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文探讨了使用生成式人工智能模型辅助硬件设计自动化的可能性,重点是高层次综合(HLS)的代码生成。研究团队发布了超过50个预训练模型和170多个编程语言数据集,并比较了Github Copilot和DeepMind的AlphaCode等模型。面临挑战包括HDL数据稀缺、预训练知识迁移难度大和生成成本高。研究通过爬取在线仓库中的HLS程序,过滤无效代码样本,使用ChatGPT生成设计描述,构建了大规模数据集。采用参数高效微调方法QLoRA对预训练模型CodeLLaMA-7B进行微调,并开发了一个两步反馈框架,包括语法检查和功能检查。结果显示,微调显著提高了语法和功能的正确性,而链式思考(CoT)提示也显著改善了输出质量。未来研究将深入探讨生成硬件设计的表现,并量化HLS与HDL的运行成本差异。
"LLM如何助力硬件设计自动化?" 挑战与解决方案" "预训练语言模型在硬件生成中的应用前景如何?"
客服
商务合作
小程序
服务号
折叠