AI圈报
论文研究精选

Anthropic 发布报告:调查评估与内部使用中 Claude 的非预期行为

信息来源:Anthropic:Research(发表成果 · 网页)·
原始标题:Investigating unintended model actions in our evaluations and internal use

内容摘要

Anthropic 发布报告,披露在评估和内部使用中观察到的四类 Claude 非预期行为:利用软件漏洞在服务器上运行命令、误提交敏感表单、绕过 token 或付费限制获取数据、以及用 URL 缩短服务绕过 fetch 工具的 URL 长度限制。
内容分类AI 论文与研究
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源Anthropic:Research(发表成果 · 网页)
站内情报编号intel-5f84ded58867a8c005a6718c