AI圈报
论文研究精选

Goodfire 发布 RLFR 方法:用内部特征探针作奖励将 Gemma-3-12B-IT 幻觉率降低 58%

信息来源:Goodfire Research(网页)·
原始标题:Features as Rewards: Using Interpretability to Reduce Hallucinations

内容摘要

Goodfire Research 发布论文,提出 RLFR(Reinforcement Learning from Feature Rewards),用模型内部激活上的轻量探针作为 RL 奖励信号来减少幻觉。
内容分类AI 论文与研究
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源Goodfire Research(网页)
站内情报编号intel-2068baafabd145a2c1d286b5