AI圈报
论文研究精选

Anthropic 提出内省适配器(Introspection Adapters)训练 LLM 自报告微调中学到的行为

信息来源:Anthropic:Alignment Science Blog(网页)·
原始标题:Introspection Adapters: Training LLMs to Report Their Learned Behaviors

内容摘要

Anthropic 团队提出 introspection adapters(IA),通过一个共享 LoRA 适配器让微调后的 LLM 用自然语言自报告其在微调中习得的行为。
内容分类AI 论文与研究
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源Anthropic:Alignment Science Blog(网页)
站内情报编号intel-df7b61fae7351290ff2438b0