节点文献

基于压缩域特征话者识别的电视节目分类检索

COMPRESSED FEATURE BASED TV PROGRAM CLASSIFICATION AND RETRIEVAL USING SPEAKER IDENTIFICATION

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吴飞庄越挺郑科刘骏伟潘云鹤

【Author】 Wu Fei, Zhuang Yueting, Zheng Ke, Liu Junwei, Pan Yunhe (Institute of Artificial Intelligence , Zhejiang University, Hangzhou 310027)

【机构】 浙江大学人工智能研究所

【摘要】 本文提出在压缩域上直接对MPEG音频信号进行分析,达到电视节目实时分析检索目的.算法分为三步:首先利用压缩域特征对音频信号进行分割,然后应用分层方法把分割出来的音频片段粗分成音乐、语音和其它三个基本类别;由于话者身份是语音信号中的重要检索线索,最后利用隐马尔可夫链实现了与文本无关的话者识别,并用识别出来的话者身份对语音信号和其相应的视频进行标注.

【Abstract】 In order to perform real-time TV program analysis and retrieval, this paper presents to directly deal with MPEG multimedia stream using compressed features. The algorithm consists of three steps: first the MPEG audio stream is segmented using compressed features; then the segmented clips are hierarchically coarse-grained classified into three basic classes, i.e. music, speech and others; since speaker identity is an important cue for multimedia retrieval, HMM is used to implement recognition of text-independent speaker, the identified speaker identity is used to label audio speech and corresponding video.

【基金】 国家自然科学基金(69803009,69733030);教育部优秀年轻教师基金;高等学校骨干教师资助计划
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2002年01期
  • 【分类号】TN943
  • 【被引频次】6
  • 【下载频次】88
节点文献中: 

本文链接的文献网络图示:

本文的引文网络