AI圈报
模型发布 / 更新精选

DeepSeek-V4.1-Flash 发布:552B MoE 多模态模型主打 KV cache 压缩

信息来源:HuggingFace Daily Papers(社区热门论文)·
原始标题:DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

内容摘要

DeepSeek 发布 DeepSeek-V4.1-Flash,一个 552B 参数的多模态 MoE 模型,支持最长 100 万 token 上下文,模型权重已在 Hugging Face 开放。
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源HuggingFace Daily Papers(社区热门论文)
站内情报编号intel-8fa8d0058cc874b5744e8910