节点文献
基于改进LDA和K-means算法的主题句聚类
Topic sentence clustering based on improved latent Dirichlet allocation and K-means algorithm
【摘要】 针对隐含狄利克雷分布(LDA)主题个数的随机选定和传统K-means算法初始聚类中心选择的随机性等缺陷,提出一种新颖启发式的主题句聚类方法。该方法利用文档集聚类簇数与拆分为句子集中隐藏的主题数目一致特点,先通过层次聚类分析出文档集聚类簇,采用最小描述长度(MDL)剪枝算法来确定最佳聚类数n个,然后将n作为隐含狄利克雷分布的主题数目的先验参数,计算n个主题所在维度上的重要句子作为初始聚类中心,最终完成隐含主题句聚类。实验结果表明改进后聚类算法克服了噪声数据的干扰,避免了主题数的经验误差,聚类结果更精确。
【Abstract】 The Latent Dirichlet Allocation( LDA) model has an random number of themes,and the traditional K-means algorithm tends to local optimal solution since the initial cluster centers are randomly selected.To sovle those problems,a novel heuristic clustering method for topic sentences was proposed.In this algorithm,it was supposed that the number of document clusters should be similar to the number of topics hidden in sentences.Firstly,the number of documents clusters was acquired through hierarchical clustering,and the best-n clusters was processed using Minimal Description Length( MDL)algorithm.Then,the number n would be the prior arguement for the topics of LDA model,and the initial center for clusters could be determined based on n representative topic sentences.The goal is to cluster topic sentences.The clustering results of the improved algorithm are more accurate with addressing noisy datas and avoiding the empirical error of topics.
【Key words】 Latent Dirichlet Allocation(LDA); K-means algorithm; Minimal Dscription Length(MDL) algorithm; sentence clustering;
- 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2016年S2期
- 【分类号】TP391.1
- 【被引频次】9
- 【下载频次】312