AI圈报
论文研究普通

混合潜在注意力让循环模型吞吐提升7.4倍

信息来源:X:Rohan Paul (@rohanpaul_ai)·
原始标题:New Virginia Tech paper shows how cache older tokens as compact vectors that attention reads directly, and a looped LLM fits 4.0 to 8.8 t...

内容摘要

Virginia Tech 论文提出 Hybrid Latent Attention,将较旧的 token 缓存为紧凑向量供注意力直接读取,使循环 LLM 在单 GPU 上可容纳的请求量提升 4.0 至 8.8 倍,且精度损失很小。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源X:Rohan Paul (@rohanpaul_ai)
站内情报编号intel-130a678a767b96c0995e8fff