AI圈报
教程 / 实战普通

The Embedding Table Was 72% of the Model

信息来源:DEV Community·

内容摘要

Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.
内容分类AI 教程与实战
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源DEV Community
站内情报编号intel-93992f042de7a0829a86c654