当前位置:首页 > 报告详情

大语言模型与交互式智能体:开放世界中的动态推理与规划.pdf

上传人: 张** 编号:155410 2024-02-15 17页 6.39MB

1、DataFunSummit#2023大语言模型与交互式智能体:开放世界中的动态推理与规划林禹臣-Allen Institute for AI-研究员Textual EnvironmentReal-world Situations for AgentsTask planning&execution interactive environmentALF Worldsample task:put A in Bmostly easy&shortaction space:very limitedScienceWorldComplex Setup-10 locations-25 action types-

2、200+object types-multiple states-random exceptions-30 high-level task typeshttps:/yuchenlin.xyz/swiftsage/FormulationBaseline methods:Reinforcement Learning DRRN:deep reinforcement relevance networkDRRN:deep reinforcement relevance networkAction State value for action i under state s at the time tKG

3、-A2C:add a dynamic graph to constrain the selection of actions+objects CALM:use a larger LM say GPT-2 to re-rank the action candidates after Q function is computed.Baseline methods:Imitation LearningBehavior Cloning with Transformer LMsBehavior Cloning with Transformer LMsDecision Transformer w/Caus

4、al Transformers Text Decision Transformers(Behavior Cloning)in the ScienceWorld paper Action History+Observations +Env t,t-1action t+1Oracle agent On training example tasks,I can search&generate golden paths for completing the tasks.Offline Training Data(in seq2seq mode)SayCanReflexionTaskTask:Your

5、task is to boil tin.For Each Timestep t:Action 1:go to kitchen you moved to kitchenAction 2:look around In this kitchen,you can see.Actiont-1:pick up metal pot metal pot in inventory nowDemoDemo:An oracle path for the task of boiling water.Kbest generations for Action tAction t:put metal pot on stov

6、eRerankingTaskTask:Your task is to boil tin.+the of previously failed trials in the last round.Action historyAction history(you moved to kitchenAction 2:look around In this kitchen,you can see.Action t-2:thinkthink:now I need to place the metal pot on a heater

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文主要探讨了在大语言模型与交互式智能体领域中的动态推理与规划。文中提到,在开放世界中,智能体需要在一个包含多个地点、动作类型、对象类型和状态的复杂环境中执行任务规划与执行。具体方法包括基于强化学习的DRRN和KG-A2C,以及基于模仿学习的Behavior Cloning和Decision Transformer。作者还介绍了一种名为SwiftSage的混合智能体框架,该框架结合了小型语言模型和大型语言模型,特别适用于具身动作的规划和地面样式。最后,作者讨论了SwiftSage的局限性和未来的研究方向,包括更复杂任务的通用性、知识提炼、真实世界的具身机器人等。
"大语言模型如何助力智能体交互?" "混合型智能体框架SwiftSage如何工作?" "AI智能体在未来机器人技术中的潜力与挑战是什么?"
客服
商务合作
小程序
服务号
折叠