模型发布 / 更新精选
DeepSeek-V4.1-Flash 发布:552B MoE 多模态模型主打 KV cache 压缩
原始标题:DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
内容摘要
DeepSeek 发布 DeepSeek-V4.1-Flash,一个 552B 参数的多模态 MoE 模型,支持最长 100 万 token 上下文,模型权重已在 Hugging Face 开放。
内容分类AI 模型发布与更新
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源HuggingFace Daily Papers(社区热门论文)
站内情报编号intel-8fa8d0058cc874b5744e8910