节点文献

广播新闻语料识别中的自动分段和分类算法

Audio Segmentation and Classification in a Broadcast News Task

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吕萍颜永红

【Author】 Lü Ping Yan Yong-hong (Zhongke Xinli Speech Lab, Institute of Acoustics, Chinese Academy of Sciences, Beijing 100080, China)

【机构】 中国科学院声学研究所中科信利实验室中国科学院声学研究所中科信利实验室 北京100080北京100080

【摘要】 该介绍了中文广播新闻语料识别任务中的自动分段和自动分类算法。提出了3阶段自动分段系统。该方法通过粗分段、精细分段和平滑3个阶段,将音频流分割为易于识别的音频段。在精细分段阶段,文中提出两种算法:动态噪声跟踪分段算法和基于单音素解码的分段算法。仿效说话人鉴别中的方法,文中提出了基于混合高斯模型的分类算法。该算法较好地解决了音频段的多类判决问题。在“新闻联播”测试数据中的实验结果表明,该文提出的自动分段和分类算法性能与手工分段分类性能几乎相当。

【Abstract】 This paper describes the work on the development of an audio segmentation and classification system applied to a broadcast news task for Chinese language. Three-phase automatic audio segmentation algorithm is provided. Audio stream is cut to audio segments (or sentences) by simply segmentation, fine segmentation and smoothing. Two different fine segmentation algorithms are given. They are dynamic noise tracking segmentation algorithm and segmentation based on mono-phone decoder algorithm respectively. Classifier based on mixture Gaussian model is used to classify audio segment into four groups: noise, music, male and female. The experiments on “Xin Wen Lian Bo” broadcast news show the performance of automatic segmentation and classification is almost equivalent to that of manual segmentation and classification.

【基金】 中国科学院百人计划(G13BR01);国家973计划(2004CB318106)资助课题
  • 【文献出处】 电子与信息学报 ,Journal of Electronics & Information Technology , 编辑部邮箱 ,2006年12期
  • 【分类号】TP391.42
  • 【被引频次】4
  • 【下载频次】160
节点文献中: 

本文链接的文献网络图示:

本文的引文网络