论文研究普通
DuMateBench 论文:Agent 框架差异可使同一模型表现相差 27.27 个百分点
原始标题:New Stanford and other top research lab paper shows that the framework around an LLM can change agen…
内容摘要
论文《DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows》提出基于 200 个真实用户会话重建任务的基准,涵盖编码、网络调研、文档处理和内容创作,并加入缺失依赖、不稳定网络和干扰文件等环境复杂度。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:Rohan Paul (@rohanpaul_ai)
站内情报编号intel-9f00e93f78cdfab50871e458