节点文献

一种音频辅助的视频分割方法研究

Video Segmentation with the Support of Audio

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙中伟张福炎

【Author】 Sun Zhongwei, Zhang Fuyan (Department of Computer Science & Technology,Nanjing University,Nanjing,210093,China)

【机构】 南京大学计算机科学与技术系南京大学计算机科学与技术系 南京210093南京210093

【摘要】 视频分割是视频结构化组织的基础 .提出一种结合音频和视觉信息的视频分割新方法 ,即先对视频作基于边变化率的初步分割 ,然后提取音频的MFCC及其差分系数特征 ,利用广义似然比 (GLR)距离对音频信息进行相似性比较 ,并检测相应的音频变化点 .在此基础上 ,应用音频分割点对初步的视频分割进行验证 ,获得具有一定语义内容的视频段 .实验结果表明 ,方法简单有效 ,与单一的基于视觉信息的分割方法相比 ,获得的视频片段语义信息更为完整 ,同时也避免了分割的过度细碎

【Abstract】 Video segmentation is the first step to the structural organization of the video. While previous research on video segmentation primarily focuses on the pictural part, this paper presents a video segmentation method which combines the audio and visual information. With the proposed method, the edge change fraction between the adjacent frames is computed, which uses the Hausdarff distance for global motion compensation. Video data are first partitioned into segments using the edge change fraction. A distance measure derived from the generalized likelihood ratio (GLR) is used for the detection of audio changing points. Regarding the audio feature extraction for the distance measure, the MFCC coefficients and its Delta coefficients are considered. The video segmentation points are then validated or discarded with the audio changing points. In contrast to the visual-based processing, the proposed method avoids the problem of a far too fine segmentation of the video with respect to its semantic meaning.

  • 【文献出处】 南京大学学报(自然科学版) ,Journal of Naijing University (Natural Sciences) , 编辑部邮箱 ,2002年02期
  • 【分类号】TP391.41
  • 【下载频次】68
节点文献中: 

本文链接的文献网络图示:

本文的引文网络