AI圈报
教程 / 实战普通

Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

信息来源:DEV Community·

内容摘要

A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.
内容分类AI 教程与实战
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源DEV Community
站内情报编号intel-1a99c6960c4ae4b7100b0d61