论文研究普通
TT-VidT:解耦时间轴的高效运动中心视频预训练
原始标题:TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
内容摘要
TT-VidT 通过将 DINOv3 初始化的 ViT-B/16 逐帧空间路径与紧凑 Temporal Transfer Layer 结合,用 Diff Compression 从首帧外观锚点和逐帧运动 token 重建目标帧。
内容分类AI 论文与研究
内容层级普通情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源HuggingFace Daily Papers(社区热门论文)
站内情报编号intel-84cd3df7a1fc67da8db37349