节点文献

一种扩展的基于学习的视频镜头检测方法

An Extended Learning-Based Approach for Video Shot Boundary Detection

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘森方卫封化民宋国森方勇

【Author】 LIU Sen 1,FANG Wei 2,FENG Huamin 3,SONG Guosen1,FANG Yong2,31(The school of Information Science and engineering, YanSan University, Qin Huangdao066004, China)2(The school of Telecommunication Engineering, Bupt University, Beijing100876, China)3(Key Laboratory for Security and Secrecy of Information, Besti, Beijing100070, China)

【机构】 燕山大学信息工程学院北京邮电大学电信工程学院北京电子科技学院信息安全与保密重点实验室

【摘要】 视频镜头的边界检测是视频检索的基础。随着视频编辑技术的快速发展,对于各种新的编辑方法需要有新的检测方法相适应。“画中画”技术就对以往的检测方法提出了新的挑战。本文在时域多尺度边界检测(TMRA)的基础上,提出以分块色彩直方图为特征,用支持向量机(SVM)模式识别工具和滑动窗口技术,对视频帧进行分类的新方法。通过对30个小时(15,253个镜头)的新闻视频进行测试表明本方法是非常有效的,本框架基本上解决了“画中画”技术造成的错检,对渐变的检测亦有较高的准确率。

【Abstract】 Shot Boundary detection (SBD) is the basis for video retrieval. With the rapid development of video editing technology, we need new video processing techniques to cope with the new video editing techniques. One of the new challenges was posed with emerging from the sub-window technique in news video, the original method of video segmentation cannot efficiently detect the video shot boundary caused by special video technique.In this paper, we demonstrate that a temporal video sequence can be visualized as a trajectory of points in the multi-dimensional feature space. By studying different types of transitions in different resolutions, it can be observed that the shot boundary detection is a temporal multi-resolution phenomenon. From this insight, an improved temporal multi-resolution analysis (TMRA) algorithm was developed for video segmentation by using Canny wavelet and applying multi-temporal-resolution analysis on the video stream. The solution is a general framework for all kinds of video transition. We proposed Blocked Chromatic Histogram (BCH) as a feature vector, the video frame series have the temporal multi-resolution characteristics of shot presented by the wavelet transition coefficients. Traditional statistical method cannot deal with feature vectors with so high dimensions efficiently, but SVM (Supported Vector Machines) could resolve this problem. We use the SVM classifier for pattern recognition and a sliding window to dynamically classify the video frames into normal frames, gradual transition frames and CUT frames, then we clustering the classified frames into different shot categories. In one word, BCH can supply sufficient information in the video clips and SVM tools can filter the noise in the video automatically, so we can get better detection result. The testing result of the experiment on news video clips’ truth bed, which has about 30 hours (15,253 shots) in all, shows that the new method has relatively better performance for the SBD. The system is able to improve the precision of the SBD while retaining high recall. Our method can resist the influences of sub-window. So it basically resolves the difficulties of shot boundaries caused by sub-window technique in news video clips. It also greatly improves performance of locating the positions of start points and end points of GTs. We will overcome the problem by using a more wide sliding window in the future. Our system will be used to segment on the Web, and do semantic analysis combing with text around.

【关键词】 分块色彩直方图支持向量机(SVM)新闻视频
【Key words】 BCHSVMnews video
【基金】 国家自然科学基金项目[项目号:60472082]
  • 【会议录名称】 第一届建立和谐人机环境联合学术会议(HHME2005)论文集
  • 【会议名称】第一届建立和谐人机环境联合学术会议(HHME2005)
  • 【会议时间】2005-10
  • 【会议地点】中国昆明
  • 【分类号】TP391.41
  • 【主办单位】中国计算机学会、中国图象图形学学会、ACM SIGCHI中国分会、清华大学计算机科学与技术系
节点文献中: 

本文链接的文献网络图示:

本文的引文网络