当前位置:首页 >英文主页 >中英对照 > 报告详情

Kimi K1.5技术报告(英文版)(25页).pdf

上传人: Kell****reet 编号:599100 2025-02-02 25页 887.28KB

下载:

1、KIMI K1.5:SCALINGREINFORCEMENTLEARNING WITHLLMSTECHNICALREPORT OFKIMI K1.5Kimi TeamABSTRACTLanguage model pretraining with next token prediction has proved effective for scaling compute butis limited to the amount of available training data.Scaling reinforcement learning(RL)unlocks a newaxis for the

2、 continued improvement of artifi cial intelligence,with the promise that large languagemodels(LLMs)can scale their training data by learning to explore with rewards.However,priorpublished work has not produced competitive results.In light of this,we report on the training practiceof Kimi k1.5,our la

3、test multi-modal LLM trained with RL,including its RL training techniques,multi-modal data recipes,and infrastructure optimization.Long context scaling and improved policyoptimization methods are key ingredients of our approach,which establishes a simplistic,effectiveRL framework without relying on

4、more complex techniques such as Monte Carlo tree search,valuefunctions,and process reward models.Notably,our system achieves state-of-the-art reasoningperformance across multiple benchmarks and modalitiese.g.,77.5 on AIME,96.2 on MATH500,94-th percentile on Codeforces,74.9 on MathVistamatching OpenA

5、Is o1.Moreover,wepresent effective long2short methods that use long-CoT techniques to improve short-CoT models,yielding state-of-the-art short-CoT reasoning resultse.g.,60.8 on AIME,94.6 on MATH500,47.3on LiveCodeBenchoutperforming existing short-CoT models such as GPT-4o and Claude Sonnet3.5 by a l

6、arge margin(up to+550%).Kimi k1.5 long-CoTOpenAI o1OpenAI o1-miniQwQ-32B PreviewQVQ-72B-PreviewVision74.974.97171.4MathVista(Pass1)707077.370.3MMMU(Pass1)Code9494948862Codeforces(Percentile)62.562.567.253.140.6LiveCodeBench v5 24.12-25.2(Pass1)Math96.296.294.89090.6MATH 500(EM)77.577.574.463.650AIME

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文介绍了KIMI K1.5,一种使用强化学习(RL)训练的多模态语言模型。主要内容包括: 1. KIMI K1.5通过RL训练,实现了在多个基准测试和模态上的最先进推理性能,例如在AIME上达到77.5,在MATH500上达到96.2,在Codeforces上达到94百分位,在MathVista上达到74.9。 2. KIMI K1.5采用了长上下文缩放、改进的策略优化方法、简单的框架等关键技术。长上下文缩放将上下文窗口扩展到128k,改进了策略优化方法,包括在线镜像下降的变体,以及有效的采样策略、长度惩罚和数据配方优化。 3. KIMI K1.5还展示了如何将长CoT技术应用于短CoT模型,以提高其性能。例如,KIMI K1.5-short w/ rl在AIME上达到60.8,在MATH500上达到94.6,在LiveCodeBench上达到47.3,优于GPT-4o和Claude Sonnet 3.5等现有短CoT模型。 4. KIMI K1.5的RL训练系统采用了部分回放技术,有效处理了长CoT特征,实现了高效的训练和推理。 5. KIMI K1.5在文本、视觉和推理挑战中展示了强大的自然语言理解、数学、编码和逻辑推理能力。
如何通过强化学习训练大规模语言模型? 长序列强化学习如何提升模型性能? 如何将长序列模型压缩为短序列模型?
客服
商务合作
小程序
服务号
折叠