教程 / 实战普通
The Embedding Table Was 72% of the Model
内容摘要
Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.