节点文献

双流运动建模-循环一致性对齐小样本动作识别算法

Two-stream motion modeling and cycle consistency alignment for few-shot action recognition

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 胡正平董佳伟王昕宇

【Author】 HU Zhengping;DONG Jiawei;WANG Xinyu;School of Information Science and Engineering, Yanshan University;Hebei Key Lab of Information Transmission and Signal Processing, Yanshan University;

【通讯作者】 胡正平;

【机构】 燕山大学信息科学与工程学院燕山大学河北省信息传输与信号处理重点实验室

【摘要】 针对不同场景下动作时空分布不同导致视频对齐困难,进而影响视频识别准确度问题,提出对双流特征进行运动建模和循环一致性对齐的小样本动作识别方法,能够在全局帧和局部块双尺度特征建模和对齐高维运动表示。首先基于双流特征设计了运动建模框架,重塑视频序列中动作表示的时空联系,实现对视频动作的准确定位和语义性捕获;然后,为帮助模型学习动作间时空对应关系,引入循环一致性对齐机制,利用软最近邻查询的方法,高效对齐视频动作,显著改善了视频动作的错位问题;最后,结合基于注意力机制的时域交叉匹配模块,对动作类别进行推理分类。实验结果表明,该算法在SSv2、HMDB51、UCF101上分别达到68.6%、77.7%和96.9%的识别精度,实现了对视频动作的有效识别。

【Abstract】 To address the challenge of aligning videos posed by different spatio-temporal distributions of actions, which subsequently affects the accuracy of video recognition in various scenarios, a few-shot action recognition method is proposed to model and align the two-stream features through motion modeling and cycle consistency alignment. This method bases on the high-dimensional representation of motion by modeling and aligning the dual-scale features of global frames and local patch. Firstly, a motion modeling framework based on the two-stream features is designed to reshape the spatio-temporal relationship of action representations in video sequences, achieving precise localization and semantic capture of video actions. Furthermore, to facilitate the learning of spatio-temporal correspondences between actions, the cycle consistency alignment module is introduced, which can efficiently align video actions using soft nearest neighbor and significantly improve the misalignment issues. Lastly, the model combined with attention-based temporal cross matching module to infer and classify the action categories. The experimental results demonstrate this method achieves recognition accuracies of 68.6%, 77.7% and 96.9% on SSv2, HMDB51 and UCF101, respectively, and could effectively recognize video actions.

【基金】 国家自然科学基金资助项目(61771420,62001413);河北省自然科学基金资助项目(F2024203069)
  • 【文献出处】 燕山大学学报 ,Journal of Yanshan University , 编辑部邮箱 ,2025年01期
  • 【分类号】TP391.41
  • 【下载频次】16
节点文献中: 

本文链接的文献网络图示:

本文的引文网络