当前位置:首页 >英文主页 >中英对照 > 报告详情

微软:2025生成式AI红队百次测试经验白皮书(英文版)(21页).pdf

上传人: Kell****reet 编号:615126 2025-02-28 21页 1.17MB

下载:

1、Lessons from red teaming 100 generative AI products Authored by:Microsoft AI Red TeamAuthorsBlake Bullwinkel,Amanda Minnich,Shiven Chawla,Gary Lopez,Martin Pouliot,Whitney Maxwell,Joris de Gruyter,Katherine Pratt,Saphir Qi,Nina Chikanov,Roman Lutz,Raja Sekhar Rao Dheekonda,Bolor-Erdene Jagdagdorj,Eu

2、genia Kim,Justin Song,Keegan Hines,Daniel Jones,Giorgio Severi,Richard Lundeen,Sam Vaughan,Victoria Westerhoff,Pete Bryan,Ram Shankar Siva Kumar,Yonatan Zunger,Chang Kawaguchi,Mark Russinovich2Lessons from red teaming 100 generative AI productsTable of contents304Abstract07Red teaming operations09Ca

3、se study#1 Jailbreaking a vision language model to generate hazardous content12Lesson 4 Automation can help cover more of the risk landscape05Introduction08Lesson 1 Understand what the system can do and where it is applied10Lesson 3 AI red teaming is not safety benchmarking12Lesson 5 The human eleme

4、nt of AI red teaming is crucial05AI threat model ontology 08Lesson 2 You dont have to compute gradients to break an AI system 11Case study#2 Assessing how an LLM could be used to automate scams13Case study#3 Evaluating how a chatbot responds to a user in distressLessons from red teaming 100 generati

5、ve AI products14Case study#4 Probing a text-to-image generator for gender bias14Lesson 6 Responsible AI harms are pervasive but difficult to measure15Lesson 7 LLMs amplify existing security risks and introduce new ones 16Case study#5 SSRF in a video-processing GenAI application17Lesson 8 The work of

6、 securing AI systems will never be complete18ConclusionAbstractIn recent years,AI red teaming has emerged as a practice for probing the safety and security of generative AI systems.Due to the nascency of the field,there are many open questions about how red teaming operations should be conducted.Bas

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文由微软AI红队撰写,总结了他们对100多个生成式AI产品的红队测试经验。主要内容包括: 1. 介绍了微软AI红队的威胁模型本体论,用于指导红队操作。 2. 分享了八个主要经验教训,包括理解系统能力、不需要计算梯度也能破坏AI系统、AI红队不是安全基准测试、自动化可以覆盖更多风险、AI红队中人的作用至关重要、负责任的AI伤害普遍但难以衡量、大型语言模型放大了现有安全风险并引入了新的风险、保护AI系统的工作永远不会完成。 3. 提供了五个案例研究,展示了如何将本体论应用于各种安全和负责任的AI风险。 4. 讨论了AI红队中的一些误解,并提出了未来研究的开放问题。 总体来说,本文为AI红队实践提供了实用的建议,并强调了AI红队中人的作用和负责任的AI风险的重要性。
微软AI红队如何通过红队操作识别AI系统的安全风险? AI红队如何利用自动化工具PyRIT提高AI安全风险识别效率? AI红队如何评估AI系统在生成有害内容方面的风险?
客服
商务合作
小程序
服务号
折叠