节点文献

一种新的中文词自动聚类算法

A New Algorithm of Chinese Words Automatic Clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙静朱杰徐向华

【Author】 SUN Jing,ZHU Jie,XU Xianghua( Dept. of Electronic Eng., Shanghai Jiaotong Univ., Shanghai 200030, China)

【机构】 上海交通大学电子工程系上海交通大学电子工程系 上海200030上海200030上海200030

【摘要】 基于分类的统计语言模型是解决N-gram语言模型中数据稀疏问题的有效方法之一,词的自动聚类算法一直是一个难点.如何设计一种计算速度快、收敛性好的算法是关键.提出一种根据词的上下文环境,综合考虑语言模型的困惑度和词的相似度的自动聚类算法.把词的自动聚类和提高基于分类的语言模型的性能联合起来考虑.实验结果表明,该算法执行效率高、聚类效果好.

【Abstract】 Classbased statistical language model is an effective solution to the dearth of training set. It is a tough task to automatically classify words and also improtant to design a quick algorithm with good convergence. This paper proposed a method for words clustering based on the words’ context with perplexity and similarity as a measure. The algorithm combines words classification with improving the performance of classbased language model together. The algorithm is of high executing speed and good clustering performance.

【基金】 上海市科学技术委员会基础研究项目(01JC14033);美国贝尔实验室上海分部的资助项目
  • 【文献出处】 上海交通大学学报 ,Journal of Shanghai Jiaotong University , 编辑部邮箱 ,2003年S2期
  • 【分类号】TP391.12
  • 【被引频次】8
  • 【下载频次】262
节点文献中: 

本文链接的文献网络图示:

本文的引文网络