模型发布 / 更新普通
SGLang、Qwen 与 NVIDIA 在 SGLang 中实现 NVFP4 KV cache
原始标题:Blog Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache The KV cache is a fundamental building block of the modern LLM inference system. The context from multiple conversation rounds in agent sessions is cached as keys and values (KV) in GPU memory, allowi… SGLang, Qwen, and NVIDIA teams
内容摘要
SGLang、Qwen 与 NVIDIA 团队在 SGLang 中实现了 NVFP4 KV cache,用于加速长上下文与智能体推理。NVFP4 以 4 位 E2M1 值配合每 16 个值一个 FP8 block scale 和 FP32 tensor 级 scale,每 16 个值仅占 8 字节加 1 字节元数据,约为 FP8 存储的 56%。
内容分类AI 模型发布与更新
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源LMSYS:Blog(Chatbot Arena 团队)
站内情报编号intel-d63578d9f78a336ce3701161