节点文献

基于场景表示中对象特征语法分析的视频描述

Video captioning based on scene representation object features syntax analysis

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 付燕王咪咪叶鸥

【Author】 FU Yan;WANG Mi-mi;YE Ou;College of Computer Science and Technology, Xi’an University of Science and Technology;

【通讯作者】 王咪咪;

【机构】 西安科技大学计算机科学与技术学院

【摘要】 为解决基于编码器-解码器的视频描述方法中存在忽略特征语法分析,造成描述语句语法结构不清晰的问题,提出一种基于场景表示中对象特征语法分析的视频描述方法。编码阶段将视频的2D、C3D特征、对象特征和自注意力机制相结合,构建视觉场景表示模型,描述视觉特征间的依赖关系;构建视觉对象特征语法分析模型,分析对象特征在描述语句中的语法成分;解码阶段结合语法分析结果和LSTM网络模型,输出视频描述语句。所提方法在MSVD和MSR-VTT数据集进行实验,结果表明,该方法在不同评价指标方面性能较好,视频描述语句的语法结构清晰。

【Abstract】 To solve the problem of ignoring feature syntax analysis of video description method based on encoder-decoder, resulting in unclear description syntax structure, a video description method based on object feature grammar analysis in scene representation was proposed. In the coding stage, a visual scene representation model was constructed by combining 2D and C3D features of the video, as well as object features and self-attention mechanism to describe the dependence between visual features. A visual object feature grammar analysis model was constructed to analyze the grammatical components of object features in description sentences. The decoding stage combined the results of grammar analysis and the LSTM network model to output the video captioning. Experimental results on MSVD and MSR-VTT data sets show that the proposed method has good performance in different evaluation indexes, and the syntax structure of video description sentence is clear.

【基金】 陕西省自然科学基金项目(2018JQ5095);中国博士后科学基金项目(2020M673446)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2023年02期
  • 【分类号】TP391.41
  • 【下载频次】29
节点文献中: 

本文链接的文献网络图示:

本文的引文网络