节点文献

跨语言知识蒸馏的视频中文字幕生成

Cross-Lingual Knowledge Distillation for Chinese Video Captioning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 侯静怡齐雅昀吴心筱贾云得

【Author】 HOU Jing-Yi;QI Ya-Yun;WU Xin-Xiao;JIA Yun-De;Beijing Laboratory of Intelligent Information Technology,School of Computer Science,Beijing Institute of Technology;

【通讯作者】 吴心筱;

【机构】 北京理工大学计算机学院智能信息技术北京市重点实验室

【摘要】 视频字幕生成(video captioning)在视频推荐、辅助视觉、人机交互等领域具有广泛的应用前景.目前已有大量的视频英文字幕生成方法和数据,通过机器翻译视频英文字幕可以实现视频中文字幕的生成.然而,中西方文化差异和机器翻译算法性能都会影响中文字幕生成的质量.为此,本文提出了一种跨语言知识蒸馏的视频中文字幕生成方法.该方法不仅可以根据视频内容直接生成中文语句,还充分利用了易于获取的视频英文字幕作为特权信息(privileged information)指导视频中文字幕的生成.由于同一视频的英文字幕与中文字幕之间存在语义关联关系,本文方法从中学习到与视频内容相关的跨语言知识,并利用知识蒸馏将英文字幕包含的高层语义信息融入中文字幕生成.同时,通过端到端的训练方式确保模型训练目标与视频中文字幕生成任务目标的一致性,有效提升中文字幕生成性能.此外,本文还对视频英文字幕数据集MSVD扩展,给出了中英文视频字幕数据集MSVD-CN.

【Abstract】 Video captioning aims to automatically generate the natural language descriptions of a video,which requires understanding the visual content and describing it with grammatically accurate sentences.Video captioning has wide applications in video recommendation,vision assistance,human-robot interaction and many other fields,and has attracted growing attention in the fields of computer vision and natural language processing.Although remarkable progress has been made on English video captioning,using other languages such as Chinese to describe a video remains under-explored.In this paper,we investigate Chinese video captioning.However,the insufficiency of paired videos and Chinese captions makes it difficult to train a powerful model for Chinese video captioning.Since there exist many English video captioning methods and training data,it is a feasible method to perform Chinese video captioning by translating the English captions via machine translation.However,the difference between Chinese and Western cultures and the performance of machine translation algorithms will both affect the quality of generated Chinese captions.To this end,we propose a cross-lingual knowledge distillation method for Chinese video captioning.Based on a two-branches structure,our method does not only directly generate Chinese captions according to the video content,but also takes full advantage of the easily accessible English video captions as the privileged information to guide the generation of Chinese video captions.Since the Chinese and English captions are semantically correlated with respect to the video content,our method learns cross-lingual knowledge from them and utilizes knowledge distillation to integrate the high-level semantic information in English captions into Chinese captions generation.Meanwhile,the consistency between the training target and the captioning target is guaranteed by the end-to-end training strategy,thus effectively improving the performance of Chinese video captioning.Benefit from the mechanism of knowledge distillation,our method only utilizes English captions data during the training stage;and after training it can directly generate Chinese captions from the input video.To verify the universality and flexibility of our cross-lingual knowledge distillation method,we use four mainstream visual captioning models for evaluation,covering the CNN-RNN structure,RNN-RNN structure,CNN-CNN structure and model based on Top-Down attention mechanism.These models are widely used as the backbone models in a large number of visual captioning methods.Moreover,we extend the English video captioning dataset MSVD into a cross-lingual video captioning dataset with Chinese captions,called MSVD-CN.MSVD-CN contains 1970 video clips collected from the Internet and11 758 Chinese captions besides the original 41 English captions per video in MSVD.In order to reduce the annotation mistakes caused by annotators’ typos or misunderstandings of the video contents,we propose two automatic inspection methods to perform semantic and syntactic checks,respectively,on the collected manual annotations in the data collection stage.Extensive experiments are carried out on the MSVD-CN dataset,via four widely used evaluation metrics for video captioning including BLEU,METEOR,ROUGE-L,and CIDEr.The results demonstrate that the superiority of proposed cross-lingual knowledge distillation on Chinese video captioning.Furthermore,we also report some qualitative experiment results to show the effectiveness of our method.

【基金】 国家自然科学基金(62072041)资助~~
  • 【文献出处】 计算机学报 ,Chinese Journal of Computers , 编辑部邮箱 ,2021年09期
  • 【分类号】TP391.41;TP391.2
  • 【被引频次】5
  • 【下载频次】349
节点文献中: 

本文链接的文献网络图示:

本文的引文网络