教程 / 实战普通
Run GLM-5.3 Locally: Real Quant Sizes, the llama.cpp Surprise, and the reasoning_effort Trap
内容摘要
Flash on 27 August, the flagship on 28 August. Real quant sizes, why the bigger model has better tooling support than the smaller one, and the default that silently makes it feel slow.