论文研究普通
基准测试隐藏测试过严问题
原始标题:it's basically impossible to interpret evals by looking at just at the pass/fail scores these days …
内容摘要
如今仅凭通过/失败分数基本无法解读评测结果 我在基准测试中看到的许多失败,都源于过于严格的隐藏测试,某些情况下模型的答案比预期的评测结果更合理
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:Thariq (@trq212)
站内情报编号intel-1975c1d4e677ebfe300b853e