AIQB
TutorialsOrdinary

VRAM for local LLMs: why memory bandwidth sets your tokens per second

Source: DEV Community·

Summary

VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.
TierOrdinary
Published
Indexed by AIQB
SourceDEV Community
AIQB record IDintel-172054ed3676fc25038c933c