节点文献
基于LDA高频词扩展的中文短文本分类
A New Method of Key words Extraction for Chinese Short-text Classification
【摘要】 针对短文本特征稀疏、噪声大等特点,提出一种基于LDA高频词扩展的方法,通过抽取每个类别的高频词作为向量空间模型的特征空间,用TF-IDF方法将短文本表示成向量,再利用LDA得到每个文本的隐主题特征,将概率大于某一阈值的隐主题对应的高频词扩展到文本中,以降低短文本的噪声和稀疏性影响。实验证明,这种方法的分类性能高于常规分类方法。
【Abstract】 Short texts are different from traditional documents in their shortness and sparseness.Feature extension can ease the problem of high sparse in the vector space model,but feature extension inevitably introduces noise.To resolve the problem,this paper proposes a high-frequency words expansion method based on LDA.By extracting high-frequency words from each category as the feature space,using LDA to derive latent topics from the corpus,it extends the topic words into the short-text.Extensive experiments conducted on Chinese short messages and news titles show that the new method proposed for Chinese short-text classification can obtain a higher classification performance comparing with the conventional classification methods.
【Key words】 Short-text classification High frequency words LDA Feature expansion;
- 【文献出处】 现代图书情报技术 ,New Technology of Library and Information Service , 编辑部邮箱 ,2013年06期
- 【分类号】TP391.1
- 【被引频次】70
- 【下载频次】1104