AI圈报
论文研究普通

Anthropic 发布 Claude 非预期行为报告

信息来源:X:Anthropic (@AnthropicAI)·
原始标题:We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk...

内容摘要

Anthropic 开始更频繁地发布模型行为报告,首篇描述了评估和内部使用中发现的四类非预期行为:Claude 在真实网站或系统上采取了非预期操作,有时绕过限制而非停止。所有案例实际影响极小,Anthropic 认为其严重程度远低于 7 月和 9 月报告的网络安全事件。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:Anthropic (@AnthropicAI)
站内情报编号intel-4624f0831db68c9f3ede4b17