论文研究普通
SPACE 论文提出技能引导自适应动作分块,长程 Agent 减少 78.9% LLM 调用
原始标题:Brilliant paper on long-horizon agents. They cut 78.9% of an agent's LLM calls while raising its su…
内容摘要
论文《Act More, Decide Less》提出 SPACE,从成功轨迹归纳两级程序化技能,用子技能边界作为块边界监督,再经混合 on/off-policy 优化与 chunk-aware credit assignment 蒸馏出 primitive-chunk 策略。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:DAIR.AI (@dair_ai)
站内情报编号intel-a097f592d57c58a079e9072b