当前位置:首页 > 报告详情

中国人民大学:2024年迈向可解释和可理解的多模态大规模语言模型(英文版)(33页).pdf

上传人: AG 编号:606066 2024-12-01 33页 1.10MB

1、JOURNAL OF LATEX CLASS FILES,VOL.14,NO.8,OCTOBER 20241Towards Explainable and Interpretable MultimodalLarge Language Models:A Comprehensive SurveyYunkai Dang1,*Kaichen Huang1,*Jiahao Huo1,*Yibo Yan1,2Sirui Huang1Dongrui Liu3Mengxi Gao1Jie Zhang3Chen Qian3Kun Wang4Yong Liu5Jing Shao3Hui Xiong1,2Xumin

2、g Hu1,21The Hong Kong University of Science and Technology(Guangzhou)2The Hong Kong University of Science and Technology3Shanghai AI Laboratory4Nanyang Technological University5Renmin University of ChinaAbstractThe rapid development of Artificial Intelligence(AI)has revolutionized numerous fields,wi

3、th large languagemodels(LLMs)and computer vision(CV)systems drivingadvancements in natural language understanding and visualprocessing,respectively.The convergence of these technologieshas catalyzed the rise of multimodal AI,enabling richer,cross-modal understanding that spans text,vision,audio,and

4、videomodalities.Multimodal large language models(MLLMs),in par-ticular,have emerged as a powerful framework,demonstratingimpressive capabilities in tasks like image-text generation,visualquestion answering,and cross-modal retrieval.Despite theseadvancements,the complexity and scale of MLLMs introduc

5、e sig-nificant challenges in interpretability and explainability,essentialfor establishing transparency,trustworthiness,and reliability inhigh-stakes applications.This paper provides a comprehensivesurvey on the interpretability and explainability of MLLMs,proposing a novel framework that categorize

6、s existing researchacross three perspectives:(I)Data,(II)Model,(III)Training&Inference.We systematically analyze interpretability fromtoken-level to embedding-level representations,assess approachesrelated to both architecture analysis and design,and exploretraining and inference strategies that enh

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要内容概括如下: 1. 文章首先介绍了多模态大语言模型(MLLM)的发展现状,以及其在图像文本生成、视觉问答和跨模态检索等任务中的强大能力。 2. 然而,随着MLLM的复杂性和规模的增加,模型的可解释性和可解释性成为了一个关键挑战。文章指出,可解释性和可解释性对于确保透明度、可信度和可靠性至关重要。 3. 文章提出了一个新颖的框架,将现有的研究分为三个视角:数据、模型、训练和推理。 4. 在数据视角下,文章探讨了如何通过分析输入和输出数据来增强模型的可解释性。 5. 在模型视角下,文章深入分析了从标记级到嵌入级的表示,评估了与架构分析和设计相关的各种方法,并探索了增强透明度的训练和推理策略。 6. 最后,文章总结了各种方法的优势和局限性,并提出了未来研究的发展方向。 7. 文章为推进MLLM的可解释性和透明度提供了基础资源,为研究人员和实践者提供了开发更可解释和稳健的多模态AI系统的指导。
如何提高多模态大语言模型的可解释性和可解释性? 多模态大语言模型在哪些领域具有应用潜力? 如何通过数据、模型和训练推理三个维度来提高多模态大语言模型的可解释性?
客服
商务合作
小程序
服务号
折叠