当前位置:首页 >英文主页 >中英对照 > 中译版报告详情

世界数字技术院(WDTA):2024大语言模型安全性测试方法(中译版)(22页).pdf

上传人: Kell****reet 编号:162007 2024-05-15 22页 1.23MB

下载:

1、World Digital Technology Academy(WDTA)Large Language Model SecurityTesting MethodWorld Digital Technology Academy StandardWDTA AI-STR-02Edition:2024-04 WDTA 2024 All rights reserved.The World Digital Technology Standard WDTA AI-STR-02 is designated as a WDTAnorm.This document is the property of the

2、World Digital Technology Academy(WDTA)and isprotected by international copyright laws.Any use of this document,including reproduction,modification,distribution,or republication,without the prior written permission of WDTA,isprohibited.WDTA is not liable for any errors or omissions in this document.D

3、iscover more WDTA standard and related publications at https:/wdtacademy.org/.Version History*Standard IDVersionDateChangesWDTA AI-STR-021.02024-04Initial ReleaseForewordThe Large Language Model Security Testing Method,developed and issued by the World DigitalTechnology Academy(WDTA),represents a cr

4、ucial advancement in our ongoing commitment toensuring the responsible and secure use of artificial intelligence technologies.As AI systems,particularly large language models,continue to become increasingly integral to various aspects ofsociety,the need for a comprehensive standard to address their

5、security challenges becomesparamount.This standard,an integral part of WDTAs AI STR(Safety,Trust,Responsibility)program,is specifically designed to tackle the complexities inherent in large language models and providerigorous evaluation metrics and procedures to test their resilience against adversa

6、rial attacks.This standard document provides a framework for evaluating the resilience of large language models(LLMs)against adversarial attacks.The framework applies to the testing and validation of LLMsacross various attack classifications,including L1 Random,L2 Blind-Box,L3 Black-Box,and L4White-

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要介绍了世界数字技术学院(WDTA)发布的大型语言模型安全测试方法(WDTA AI-STR-02)。该标准旨在评估大型语言模型(LLM)在面对对抗性攻击时的抵御能力,以提高使用LLM构建的AI系统的安全性和可靠性。 主要内容包括: 1. 定义了LLM、对抗样本、对抗攻击等关键术语。 2. 将LLM的对抗攻击分为四类:L1随机攻击、L2盲盒攻击、L3黑盒攻击和L4白盒攻击。 3. 提出了评估LLM对抗攻击的指标,包括攻击成功率(R)和下降率(D)。 4. 提供了LLM对抗攻击测试的最小样本量和测试程序。 5. 附录A提供了LLM对抗攻击风险的参考列表。 总体而言,该标准为评估LLM在对抗性攻击下的安全性提供了一个全面的框架,有助于开发者和组织识别和缓解潜在的安全漏洞,从而提高AI系统的安全性和可靠性。
"如何评估大型语言模型的安全性?" "大型语言模型如何抵御对抗性攻击?" "如何测试大型语言模型对风险的响应能力?"
客服
商务合作
小程序
服务号
折叠