论文研究普通
清华等提出 One-Shot OPD:单条查询达全量数据 87% 增益
原始标题:Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix...
内容摘要
清华 NLP(OpenBMB 成员)联合中科院大学、东北大学、UIUC 和约翰霍普金斯大学提出 One-Shot OPD,将 OPD 训练集压缩到一条查询:数学任务上从 59.1 提升至 68.5(300 步),达到全量数据 OPD 69.8 的 87% 增益,并在代码、指令跟随和智能体工具使用上跨 Qwen、Llama、OLMo 成立。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:面壁智能 OpenBMB (@OpenBMB)
站内情报编号intel-440c4e561f309adcdd578a77