节点文献

基于MPEG-7视听特征的恐怖暴力视频检测

Detection of Violent Video with Audio-video Features Based on MPEG-7

【作者】 刘伟

【导师】 彭玉青;

【作者基本信息】 河北工业大学 , 计算机应用技术, 2014, 硕士

【摘要】 随着多媒体与网络技术的不断发展,网络上视频种类繁多,其中不乏大量影响青少年儿童身心健康的恐怖暴力类视频。然而,目前却缺乏有效的自动检测和过滤手段。传统的人工标注方法与基于单模态的视频检测方法已经不能适用于这个信息化社会,基于视听双模态特征的视频自动检测技术逐渐引起人们越来越多的关注,它为对日益增多的视频进行管理提供了便利。好的视频检测系统主要取决于视频特征的提取和检测算法的选取两个方面。本课题在特征提取方面,利用MPEG-7音频和视觉描述子进行恐怖暴力视频检测,通过对大量MPEG-7描述子的分析研究,根据待分类视频的特征,有针对性的提取了相应的音频、颜色、纹理、空间、时间、运动等特征,并对部分MPEG-7特征进行了补充和完善:增加了音频瞬时时间特性,得到对音色信息更加完备的表征;提出了新的利用权值计算视频主颜色的方法;自定义了视频运动强度描述子。在视频检测方法方面,采用基于BP网络的视频检测模型,并用遗传算法优化BP网络的初始权值和阈值,利用该模型融合音视觉特征,并对视频类别进行检测。采用上述方法,本文对恐怖暴力、音乐、动画、新闻四类视频进行了检测,取得了较高的查全率和查准率。实验结果表明,本文通过分析选择的特征具有代表性和区分性,即能有效全面表征视频特征,突出不同类别视频的差异,又不会因特征选择的盲目性而导致特征维数过高,降低了数据量;遗传算法优化的神经网络的融合模型提高了系统的鲁棒性;采用融合音视觉特征的方法与单模态特征相比明显提高了视频检测的效果。

【Abstract】 With the development of multimedia and network, there are a wide variety of video onthe network, in which many violent videos exist, and they affect children’s physical andmental health, but there isn’t an effective way to detect the violent videos. Traditionalmanual annotation method and video detection method using single-mode are not applied inthis information society. More and more people are concerned about automatic detection ofvideo based on audio-visual feature. It applies a convenient way for managing the growingnumber of videos.Good video detection system mainly depends on video feature extraction and detectionalgorithm. In feature extraction, the passage uses MPEG-7audio and visual descriptors todetect violent video. Through the studying of many MPEG-7descriptors, the new methodtargeted chosen the features about audio, time, color, texture, space, motion. Parts ofMPEG-7descriptors were added and improved: instantaneous feature of audio was added,motion intensity descriptor was customized, and a new method to extract dominant color ofvideo was proposed. In detection algorithm, BP neural network optimized by geneticalgorithms is used to fuse the audio and visual features, and detect the video category.Using the above method, this paper detects four kinds of videos which are violentvideo, cartoon, music and news. And higher recall and precision are showed. Experimentshows that these selected features are representative, discriminative and not only can reducethe data redundancy, but also describe audio information effectively and comprehensively.Fusion model of neural network is more robust. And the method of fusing audio and visualfeatures improves the effect of video detecting obviously.

节点文献中: