节点文献
基于多模注意力机制的密集型视频描述
Dense video captioning research based on multi-mode attention mechanism
【摘要】 为了解决密集型视频描述(dense video captioning, DVC)任务中视频特征利用不充分,视频定位分段不准确,语义描述效果不丰富的问题,采用多模注意力机制的密集型视频描述方法,提取视频中的视觉特征、音频特征和语音特征.通过多模注意力机制,在编码器中计算不同模态视频帧特征间的关联程度,在解码器中计算描述词序列特征与编码器输出的多模态视频帧特征间的关联程度,并将编码器、解码器输出特征分别作用于视频定位分段模型和语义描述模型获得视频分段和分段描述.提出的方法在ActivityNet Captions数据集上进行了理论分析和实验验证,其中F1-score达到60.09,METEOR指标达到8.78.该方法有效提高了视频定位分段和语义描述的准确性.
【Abstract】 In order to solve problems by the example of insufficient utilization of video features, inaccurate video positioning and segmentation, and scarce semantic captioning effect in the Dense Video Captioning(DVC) task, a dense video captioning method based on multi-modal attention mechanism is adopted. The visual features, audio features and speech features in the video are extracted, and the attention mechanism is introduced. The correlation degree between video frame features of different modes is calculated through the multi-mode attention mechanism in the encoder. The correlation degree between the captioning word sequence features and the multi-modal video frame features output by the encoder is calculated through the multi-mode attention mechanism in the decoder. The output features of encoder and decoder are applied to video segmentation module and semantic captioning module respectively to obtain video segmentation and segmentation captioning. The method has been through theoretical analysis and experimental verification on ActivityNet Captions data set. The F1-score index reaches 60.09 and the METEOR index reaches 8.78,which has improved the accuracy of video locating segmentation and semantic captioning effectively.
【Key words】 dense video captioning; multi-mode video features; feature fusion; multi-mode attention mechanism;
- 【文献出处】 曲阜师范大学学报(自然科学版) ,Journal of Qufu Normal University(Natural Science) , 编辑部邮箱 ,2023年02期
- 【分类号】TP391.41
- 【下载频次】19