节点文献

基于LDA模型的文本聚类研究

Document Clustering Method Based on LDA Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 董婧灵李芳何婷婷涂新辉万剑

【Author】 Dong Jing-ling~(1,2),Li Fang~(1,2),He Ting-ting~(1,2),Tu Xin-hui~(1,2),Wan Jian~(1,2) 1 Department of Computer Science,HuaZhong Normal University,Wuhan 430079 2 Network Media Branch,National Language Resources Monitoring and Research Center,Wuhan 430079

【机构】 华中师范大学计算机科学与技术系国家语言资源监测与研究中心网络媒体语言分中心

【摘要】 LDA(Latent Dirichlet Allocation)是近年来提出的一种具有文本主题表示能力的非监督学习模型。本文提出了一种基于LDA主题模型的文本聚类和聚簇描述方法。利用LDA模型挖掘隐藏在文本内的不同主题与词之间的关系,得到文本的主题分布;并将此分布作为特征融入到传统的向量空间模型来计算相似度进而对文本进行聚类;再利用主题信息对聚类结果进行聚簇描述。实验结果表明本文的方法能够明显地提高聚类的效果。

【Abstract】 Latent Dirichlet Allocation(LDA) is an unsupervised model which exhibits superiority on latent topic modeling of text data in the research of recent years.This paper presents a method which improves effectiveness of text clustering by using LDA model.It can mine the hidden relationship between the different topics and the words from texts,and get the topic distribution.Then we mix the distribution as feature into the traditional VSM,and use the topics to describe the clustering result Experimental results show that the method can improve clustering quality effectively.

【关键词】 主题模型LDA文本聚类
【Key words】 topic modelLatent Dirichlet Allocationtext clustering
【基金】 国家自然科学基金重大研究计划课题(90920005);国家自然科学基金项目(61003192);973国家重点基础研究发展计划课题(2007CB310804);教育部哲学社会科学研究重大课题攻关项目(08JZD0032);教育部/国家外国专家局高等学校学科创新引智计划课题(1307042);湖北省自然科学基金计划项目(2009CDB145);武汉市晨光计划项目(201050231067);华中师范大学中央高校基本科研业务费项目(CCNU10A02009,CCNU10C01005)
  • 【会议录名称】 中国计算语言学研究前沿进展(2009-2011)
  • 【会议名称】第十一届全国计算语言学学术会议
  • 【会议时间】2011-08-20
  • 【会议地点】中国河南洛阳
  • 【分类号】TP391.1
  • 【主办单位】中国中文信息学会
节点文献中: