当前位置:首页 > 报告详情

大型代码语言模型:探索现状、机遇与挑战.pdf

上传人: 竿*** 编号:981523 2025-11-29 57页 6.91MB

1、Large Language Models for CodeLoubna Ben Allal,Machine Learning Engineer,Science teamAbout me-ML Engineer Hugging Face-Graduated from ENS Paris Saclay&Ecole des Mines de Nancy-Working on LLMs for code&Synthetic data:“The Stack,StarCoder,Cosmopedia.”LoubnaBenAllal1https:/loubnabnl.github.io/How it st

2、arted:GitHub Copilot in 2021ML+Code=Productivity https:/ Engines+ML lead to 6%reduction in code iterations 3%of code generated by model ButAPI:Model:XData:XCode:XHow its going:Over 1.7k open models trained on codeHow did we get here?Strong Instruction-tuned and base modelsHow are code LLMs trained?W

3、hat you need to train(code)LLMs from scratchTransformer ModelUntrained ModelPretrained“Base”ModelSupervised Finetuned(SFT)ModelRLHFChat LLM(e.g.GPT-4)Training Generative AI Models Untrained ModelPretrained“Base”ModelSupervised Finetuned(SFT)ModelRLHFChat LLM(e.g.GPT-4)Training Code LLMsInstruction d

4、ataset for code:“write a function”“solve a bug”.The Landscape of code LLMs The Stack dataset StarCoder StarCoder2 3B,7B,15B sizesStarChat2(with H4 team)DeepSeek-Coder1B,7B,33BDeepSeek-Coder-InstructCodeLlama 7B,13B,70BCodeLlama-InstructOthers:StableCode from StabilityAI,CodeGen from SalesForce&LLMs

5、like Mixtral,DBRX,Qwen&YiBigCode:open-scientific collaborationWe are building LLMs for code in a collaborative way:-Full data transparency-Open source processing and training code-Model weights released with commercial friendly license1100+researchers,engineers,lawyers,and policy makersClosed Source

6、 Training data&sources not disclosedModel weights not public Sending data to external APIsNot reproducibleClosed Source Training data and sources not disclosedModel weights not public Sending data to external APIsNot reproducibleOpen Source Public data with inspection and opt-out toolsModel weights

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
根据报告的内容,全文主要内容概括如下: - **代码大型语言模型(LLMs)发展迅速**:自2021年GitHub Copilot推出以来,代码LLMs在代码迭代减少和代码生成效率上取得显著成果。 - **训练方法**:基于预训练模型,通过监督学习和指令微调进行训练。 - **主要模型和数据集**:包括Stack、StarCoder、StarChat2、DeepSeek-Coder等,模型规模从1B到70B不等。 - **开源与闭源**:开源模型如Stack、StarCoder等提供数据透明度和模型权重,而闭源模型则不公开数据和权重。 - **定制化**:通过数据预处理、指令微调、工具使用等方法定制化LLMs。 - **评估**:使用标准基准或定制基准评估模型性能。 - **部署**:通过Hugging Face Inference endpoints等平台部署模型。 - **未来方向**:包括构建更好的开源模型、数据透明度和治理、评估与推理、以及LLM系统的发展。 核心数据: - 代码迭代减少6% - 模型生成代码占比3% - Stack模型规模:15B参数 - StarCoder模型规模:8096个token的上下文长度 - 训练时间:24天
"代码LLM如何训练?" "开源代码LLM有哪些?" "代码LLM的未来方向?"
客服
商务合作
小程序
服务号
折叠