论文研究精选
Ling-3.0-flash 在 4 块 Blackwell GPU 上如何将批处理 1 解码延迟降低 54%
原始标题:Blog Chasing the Batch-1 Floor: Ling-3.0-flash Speculative Decode on Blackwell Batch-1 decode keeps getting more important. Xiaomi MiMo, for example, announced MiMo-V2.5-Pro UltraSpeed in June, claiming 1,000 tok/s decode on a one-trillion-parameter MoE model. Batch 1 gives an … RadixArk SGLang Team, Ant Ling Infra Team
内容摘要
蚂蚁 Ling Infra 团队与 RadixArk SGLang 团队将 Ling-3.0-flash 混合线性注意力 MoE 模型的单请求解码速度从 288 tok/s 提升至 606 tok/s,平均 TPOT 从 3.33 ms 降至 1.53 ms。
内容分类AI 论文与研究
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源LMSYS:Blog(Chatbot Arena 团队)
站内情报编号intel-e6aca8117c07f21e35cc0a21