论文研究精选
Anthropic 提出内省适配器(Introspection Adapters)训练 LLM 自报告微调中学到的行为
原始标题:Introspection Adapters: Training LLMs to Report Their Learned Behaviors
内容摘要
Anthropic 团队提出 introspection adapters(IA),通过一个共享 LoRA 适配器让微调后的 LLM 用自然语言自报告其在微调中习得的行为。
内容分类AI 论文与研究
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源Anthropic:Alignment Science Blog(网页)
站内情报编号intel-df7b61fae7351290ff2438b0