节点文献

基于语言-视觉对比学习的多模态视频行为识别方法

Multi-modal Video Action Recognition Method Based on Language-visual Contrastive Learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张颖张冰冰董微安峰民张建新张强

【Author】 ZHANG Ying;ZHANG Bing-Bing;DONG Wei;AN Feng-Min;ZHANG Jian-Xin;ZHANG Qiang;School of Computer Science and Engineering, Dalian Minzu University;Institute of Machine Intelligence and Bio-computing, Dalian Minzu University;Faculty of Electronic Information and Electrical Engineering, Dalian University of Technology;

【通讯作者】 张建新;

【机构】 大连民族大学计算机科学与工程学院大连民族大学机器智能与生物计算研究所大连理工大学电子信息与电气工程学部

【摘要】 以对比语言-图像预训练(Contrastive language-image pre-training, CLIP)模型为基础,提出一种面向视频行为识别的多模态模型,该模型从视觉编码器的时序建模和行为类别语言描述的提示学习两个方面对CLIP模型进行拓展,可更好地学习多模态视频表达.具体地,在视觉编码器中设计虚拟帧交互模块(Virtual-frame interaction module, VIM),首先,由视频采样帧的类别分词做线性变换得到虚拟帧分词;然后,对其进行基于时序卷积和虚拟帧分词移位的时序建模操作,有效建模视频中的时空变化信息;最后,在语言分支上设计视觉强化提示模块(Visual-reinforcement prompt module,VPM),通过注意力机制融合视觉编码器末端输出的类别分词和视觉分词所带有的视觉信息来获得经过视觉信息强化的语言表达.在4个公开视频数据集上的全监督实验和2个视频数据集上的小样本、零样本实验结果,验证了该多模态模型的有效性和泛化性.

【Abstract】 This paper presents a novel multi-modal model for video action recognition, which is built upon the contrastive language-image pre-training(CLIP) model. The presented model extends the CLIP model in two ways, i.e.,incorporating temporal modeling in the visual encoder and leveraging prompt learning for language descriptions of action classes, to better learn multi-modal video representations. Specifically, we design a virtual-frame interaction module(VIM) within the visual encoder that transforms class tokens of sampled video frames into virtual-frame tokens through linear transformation, and then temporal modeling operations based on temporal convolution and virtual-frame token shift are performed to effectively model the spatio-temporal change information in the video. In the language branch, we propose a visual-reinforcement prompt module(VPM) that leverages an attention mechanism to fuse the visual information, carried by the class token and visual token which are both output by the visual encoder, to enhance the language representations. Fully-supervised experiments conducted on four publicly available video datasets, as well as few-shot and zero-shot experiments conducted on two video datasets, demonstrate the effectiveness and generalization capabilities of the proposed multi-modal model.

【基金】 国家自然科学基金(61972062);辽宁省应用基础研究计划(2023JH2/101300191);国家民委中青年英才培养计划资助~~
  • 【文献出处】 自动化学报 ,Acta Automatica Sinica , 编辑部邮箱 ,2024年02期
  • 【分类号】TP391.41
  • 【下载频次】189
节点文献中: 

本文链接的文献网络图示:

本文的引文网络