节点文献

基于LDA高频词扩展的中文短文本分类

A New Method of Key words Extraction for Chinese Short-text Classification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 胡勇军江嘉欣常会友

【Author】 Hu Yongjun1 Jiang Jiaxin2 Chang Huiyou3 1(Business School,Sun Yat-Sen University,Guangzhou 510275,China) 2(School of Information Science and Technology,Sun Yat-Sen University,Guangzhou 510006,China) 3(School of Software,Sun Yat-Sen University,Guangzhou 510006,China)

【机构】 中山大学管理学院中山大学信息科学与技术学院中山大学软件学院

【摘要】 针对短文本特征稀疏、噪声大等特点,提出一种基于LDA高频词扩展的方法,通过抽取每个类别的高频词作为向量空间模型的特征空间,用TF-IDF方法将短文本表示成向量,再利用LDA得到每个文本的隐主题特征,将概率大于某一阈值的隐主题对应的高频词扩展到文本中,以降低短文本的噪声和稀疏性影响。实验证明,这种方法的分类性能高于常规分类方法。

【Abstract】 Short texts are different from traditional documents in their shortness and sparseness.Feature extension can ease the problem of high sparse in the vector space model,but feature extension inevitably introduces noise.To resolve the problem,this paper proposes a high-frequency words expansion method based on LDA.By extracting high-frequency words from each category as the feature space,using LDA to derive latent topics from the corpus,it extends the topic words into the short-text.Extensive experiments conducted on Chinese short messages and news titles show that the new method proposed for Chinese short-text classification can obtain a higher classification performance comparing with the conventional classification methods.

【基金】 国家863计划基金项目“农产品全供应链多源信息感知技术与产品开发”(项目编号:2012AA101701-03)的研究成果之一
  • 【文献出处】 现代图书情报技术 ,New Technology of Library and Information Service , 编辑部邮箱 ,2013年06期
  • 【分类号】TP391.1
  • 【被引频次】70
  • 【下载频次】1104
节点文献中: 

本文链接的文献网络图示:

本文的引文网络