节点文献
基于词向量特征扩展的中文短文本分类研究
CHINESE SHORT TEXT CLASSIFICATION BASED ON WORD VECTOR EXTENSION
【摘要】 针对中文短文本词汇较少、噪声多、特征稀疏的特性,为了提高短文本分类精确度,提出一种基于维基百科词向量的特征扩展算法。利用维基百科语料集训练词向量,通过对文本关键词高相似度词集进行特征扩展,并将得到的文本用传统的分类器进行分类。实验结果表明,所提方法在短文本分类精确度上要优于其他的文本特征扩展算法。
【Abstract】 In order to improve the accuracy of short text classification,because of the characteristics of less,more noise and less characteristic of Chinese short text,a feature extension algorithm based on Wikipedia word vector was proposed to improve the accuracy of the text classification. By using the Wikipedia corpus training word vectors,the characteristics of the word sets of text keywords extended,and the text was classified by traditional classifier. The experimental results demonstrate that the proposed method is better than other text feature extension algorithms.
【Key words】 Short text; Wikipedia; Feature extension; Word vector; Text classification;
- 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2018年08期
- 【分类号】TP391.1
- 【被引频次】44
- 【下载频次】539