AI圈报
模型发布 / 更新普通

SGLang、Qwen 与 NVIDIA 在 SGLang 中实现 NVFP4 KV cache

信息来源:LMSYS:Blog(Chatbot Arena 团队)·
原始标题:Blog Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache The KV cache is a fundamental building block of the modern LLM inference system. The context from multiple conversation rounds in agent sessions is cached as keys and values (KV) in GPU memory, allowi… SGLang, Qwen, and NVIDIA teams

内容摘要

SGLang、Qwen 与 NVIDIA 团队在 SGLang 中实现了 NVFP4 KV cache,用于加速长上下文与智能体推理。NVFP4 以 4 位 E2M1 值配合每 16 个值一个 FP8 block scale 和 FP32 tensor 级 scale,每 16 个值仅占 8 字节加 1 字节元数据,约为 FP8 存储的 56%。
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源LMSYS:Blog(Chatbot Arena 团队)
站内情报编号intel-d63578d9f78a336ce3701161