节点文献

基于语义的视频浏览系统中的关键技术研究

Research on the Key Technologies of Semantic Based Video Browsing System

【作者】 钱学明

【导师】 信息与通信工程;

【作者基本信息】 西安交通大学 , 刘贵忠, 2007, 博士

【摘要】 方便可得的视频媒体自然极大地丰富着人们的生活和工作,但是人们在享受信息时代方便快捷的服务时,同样也面临着一个长期困扰的问题:如何快捷地从网络视频媒体库中定位到自己真正感兴趣的内容。该问题是视频分析和检索领域中一个富有挑战性的研究课题。为了解决该问题,通常需要对视频进行语义的分类,以充分挖掘其中语义对象、概念以及他们之间的内在因果关系,以期望给用户提供类似于Internet上基于关键词的文本查询方式。另外,对视频媒体内容按照一个灵活的方式进行展示,以辅助用户进行内容的查询和浏览,就象翻书那样方便地从视频媒体库中找到自己感兴趣的内容。本文是针对基于语义的视频浏览系统中的关键技术研究而展开的,其中的工作包括如下几个方面:(1)基于压缩域信息的特征分析和特征提取。由于视频数据量大,而且视频数据不同于图像和文字等媒体内容,其最大的特点是存在极大的时间和空间上的冗余,因此视频媒体在存储和传输前都要进行去冗余的压缩编码。如何有效利用压缩域信息进行快速高效的特征分析和特征提取是我们研究的主要内容之一。在相关的研究中,使用I帧DCT系数来近似表示图像的纹理,用DC图像及其直方图来近似表达原始图像以及其直方图以快速进行分析,用压缩域中的运动矢量场来进行快速的运动特征描述等。这些为本文后续工作中采用压缩域信息的特征表示以及快速的特征分析和提取方法奠定了坚实的基础。(2)镜头边界检测以及基于语义的镜头分类。镜头边界检测也即场景切换类型检测是视频分析中的一个基本环节。本文从镜头切换的数学模型出发推导了Flashlight和Fade in/out的累计直方图差的一般特性,并用压缩域的DC图像的直方图来进行快速的检测。从累计直方图的特性,不仅可以检测出Fade in/out并且能够进行Fade in/out所对应的语义类型识别。在对Flahslight和Fade in/out确认的基础上进行Cut、Dissolve和其它类型的场景切换检测,极大地提高了镜头边界检测的性能。进行镜头的语义分类是视频检索中的一个重要环节,从语义的镜头类型信息以及相应的音频数据类型信息能够有效地进行视频故事单元的主题思想理解。本文对视频按照其所在领域知识、视频编辑中的约定俗成的制作手法,以及摄像机的运动模式将足球比赛视频中的镜头划分成一系列语义的类别。并且将音频按照时域和谱域的能量分布特性划分成纯说话、静音、纯音乐、含有背景噪声的说话以及含有背景音乐的说话片段等五种类型。这种场景切换类型识别、语义的镜头类型特征以及音频数据类型等都为我们进行后续高级语义事件和故事单元的检测和分类提供了重要的信息。(3)字幕检测、定位、跟踪、分割以及字幕类型划分方法。利用MPEG压缩域中I帧DCT系数所表达的纹理特征进行快速、高效的字幕检测和定位。并用压缩域中的DCT系数特征来对字幕的出现和消失帧予以快速的跟踪,最终融合包括前背景以及视频字幕的时间冗余特性来进行高效的字幕分割。并且对H.264/AVC和MPEG压缩域中的字幕检测性能进行了对比分析。将字幕按照其活动性以及存在时间的长度信息划分成滚动字幕、长期字幕、说话内容字幕和标题字幕等4种类型,以辅助进行基于语义的事件和故事单元检测和分类。(4)全局运动估计、基于摄像机运动模式的镜头细分和基于全局/局部运动相结合的应用。我们使用压缩域中的运动矢量场来进行快速的全局运动估计。其中包括基于运动矢量组的全局运动估计和基于遗传算法的全局运动估计。利用全局运动信息,将足球比赛视频中的全景镜头进行语义的细分。这种细分后的语义镜头类型信息为进行足球视频中的语义事件挖掘提供了重要的参考依据。另外,提出了基于GM/LM视频字幕遮挡区域恢复以及视频通信系统中的错误恢复,达到了较好的恢复效果。(5)语义事件和故事单元的检测和分类。融合语义的镜头类型信息、字幕类型信息、摄像机运动模式、视频领域相关知识来将体育比赛视频序列划分成进球、射门、犯规、定位球和普通等5种完备的事件集合。并且在此基础上对精彩事件按照摄像机运动模式分类到比赛中的两支球队中。融合视觉信息、字幕类型信息和音频类型特征来进行新闻视频故事单元的检测和分类。这种事件和故事单元的分类方式为进行基于高级语义特征的视频检索、摘要和浏览提供可能。(6)一个统一的灵活方便的视频摘要和浏览系统框架。在事件和故事单元的检测和分类基础上,按照书目编排的目录即ToC结构来有效进行新闻和体育视频内容组织。并提出了一种通用的基于ToC的视频内容浏览系统框架。在该框架中,按照事件和故事单元的分类情况,给出了一种可分级的视频摘要和浏览方案。不同级所生成的摘要,能够很好地提供对视频内容浏览的形式,使用户可以象浏览书本那样方便快捷地进行视频内容浏览,并对感兴趣的内容进行快速的定位。

【Abstract】 With the development of information science and techonology, video resources are becoming indispensable parts of people’s daily life. When we enjoy the convience of information science, a boring problem appears at the same time, namely, how to search the interested video sections from a huge amount of video databases in the internet. It is a challenging problem for researchers in video analysis, indexing and retrieval. For the flexible video searching task in the management of large scale video database, one effective way is to parse video sequences into semantic shots and to mine various semantic objects, concepts, and causalities among them according to certain rules. Browsing the content of a video sequence effectively like browsing a book is a popular way, where viewers can easily find what they want by consulting the table of content of the video databases. We have accomplised some fundamental work toward semantic based video analysis, indexing, retrieval and browsing, which are summarized as follows:(1) Feature analysis and extraction using the compressed domain related information. Video sequences are usually compressed using some standards, such as MPEG, for reducing the redundancy, which is a fundamental step for video transmission and storage. Using the compressed domain related information for feature analysis and extraction can speed up the process of video content analysis, indexing, retrieval and browsing. We employ the AC coefficients of a block to extract its texture information, and use the DC image to approximately represent the original frame. Moreover, motion vector field information of compressed video is also utilized to calculate the motion related information.(2) Scene change detection, scene change type recognition, and semantic based shot classification. Scene change detection, which is also named as shot boundary detection, is one of the fundamental problems in video analysis, indexing, and retrieval. Moreover, it is the minimum unit used in video edition and production. Editors and movie makers represent certain ideas through the connection of different shots. Starting from the mathematical models of flashlight and fade in/out, their excellent characteristics in accumulating histogram difference are deduced, which could not only help us to detect them effectively but also to recognize their types at the same time. Cut, dissolve, and other types of scene change are detected effectively by removing the well detected flashlight and fade in/out. Semantic based shot classification provides middle level semantic information for video content analysis and retrieval. The motive of a story unit or event clip can be discovered from the semantic shot type information. By integrating the domain related knowledge, video production knowledge and camera motin pattern information, a soccer video is segmented into a set of semantic shots. Moreover, audio data is classified into five categories: silence, pure speech, speech with background noise, speech with music background and pure music, according to the temporal domain and spectral domain related features. The semantic shot classification and audio data classification are fundamental steps toward the detection and classification of the high level semantic events and story units.(3) Text detection, localization, tracking, segmentation, and type classification. We use the AC coefficients of a block in the MPEG compressed domain to represent the texture information of the block, and to carry out fast text detection, localization and tracking for the starting and ending frames for each text. The foreground and background of video texts are integrated in text segmentation, which can reduce the influence of complex backgrounds effectively. The text detection performance of H.264/AVC and MPEG compressed domain information is compared with each other. Detected overlaid texts are classified into rolling text, long-term text, speech content text, and title text according to their motion activities and lifetimes. The text type information is fused in semantic based event and story unit classification, and text-oriented video abstraction.(4) Global motion estimation, camera motion based shot refinement, and global/local motion based applications. An MV group based global motin estimation method and a genetic algorithm based method using the verified motion field of a compressed video are proposed. From the estimated camera motion information, the global shot of soccer video is further classified into several semantic categories. Moreover, a GM/LM based text occluded region recovery and error concealment in video transmission systems are proposed, which improves the recovering results effectively compared to the GM or LM method alone.(5) Semantic based event and story unit detection and classification. The boundaries of an individual event clip and story unit are determined adaptively according to the domain knowledge and production effect before recognizing their types. Semantic shot type information, text type information of an event clip, domain related knowledge, and video production knowledge is integrated for soccer video event detection and classification. Each event clip is classified into one of the five types: shoots, goals, fouls, placed kicks and normal kicks. Moreover, the highlight events are further classified into each team according to the dominant camera motion pattern information. And for a news video, each story unit is classified into one of the 9 categories according to the multi-modal audio-visual clues. The semantic based event and story unit classification results provide a possible way for minimizing the semantic gap in video indexing and retrieval.(6) A unified video indexing, retrieval, and browsing framework is proposed similar to the ToC of a book, based on the event and story unit classification results. It is a heuristic framework, which provides a three layer video abstraction structure. Viewers can browse the video content like reading books, navigate in the video content freely and localize the interested segments effectively.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络