教程 / 实战普通
My factual-recall tasks were scoring format, not facts
内容摘要
I built a detector that decides whether a day's drift run contains anything worth writing up. The first thing it did was accuse my own suite. One task, `fact-element`, had failed o