节点文献
一种新的中文词自动聚类算法
A New Algorithm of Chinese Words Automatic Clustering
【摘要】 基于分类的统计语言模型是解决N-gram语言模型中数据稀疏问题的有效方法之一,词的自动聚类算法一直是一个难点.如何设计一种计算速度快、收敛性好的算法是关键.提出一种根据词的上下文环境,综合考虑语言模型的困惑度和词的相似度的自动聚类算法.把词的自动聚类和提高基于分类的语言模型的性能联合起来考虑.实验结果表明,该算法执行效率高、聚类效果好.
【Abstract】 Classbased statistical language model is an effective solution to the dearth of training set. It is a tough task to automatically classify words and also improtant to design a quick algorithm with good convergence. This paper proposed a method for words clustering based on the words’ context with perplexity and similarity as a measure. The algorithm combines words classification with improving the performance of classbased language model together. The algorithm is of high executing speed and good clustering performance.
【关键词】 自动聚类;
分类语言模型;
困惑度;
相似度;
算法;
【Key words】 words automatic clustering; class language model; perplexity; similarity; algorithms;
【Key words】 words automatic clustering; class language model; perplexity; similarity; algorithms;
【基金】 上海市科学技术委员会基础研究项目(01JC14033);美国贝尔实验室上海分部的资助项目
- 【文献出处】 上海交通大学学报 ,Journal of Shanghai Jiaotong University , 编辑部邮箱 ,2003年S2期
- 【分类号】TP391.12
- 【被引频次】8
- 【下载频次】262