教程 / 实战普通
My Eval Passed Because the Model Had Already Seen the Answers
内容摘要
My first green eval run was worthless: the training data and the test set came from the same sources, so the model was reciting. Fixing that exposed two more traps. Three runs of one unchanged model scored 31, 21 and 31 out of 50, and a later set that looked steady at 27, 28, 27 was hiding 29 of 50 tests changing verdict underneath.