AIQB
TutorialsOrdinary

The Embedding Table Was 72% of the Model

Source: DEV Community·

Summary

Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.
TierOrdinary
Published
Indexed by AIQB
SourceDEV Community
AIQB record IDintel-93992f042de7a0829a86c654