AI圈报
产品发布 / 更新精选

Unsloth 让 Qwen3.8-Flash 本地推理提速 1.7 倍

信息来源:X:Unsloth (@UnslothAI)·
原始标题:Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO…

内容摘要

Unsloth 通过 MTP(Multi-Token Prediction)让 Qwen3.8-Flash-Next 本地推理提速约 1.3 至 1.7 倍且精度不变,GGUF 在单张 RTX PRO 6000 上可达 170 tokens/s(基线 100 tokens/s)。
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:Unsloth (@UnslothAI)
站内情报编号intel-c1514554d6a918252ad72d35