节点文献

基于词向量特征扩展的中文短文本分类研究

CHINESE SHORT TEXT CLASSIFICATION BASED ON WORD VECTOR EXTENSION

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 雷朔刘旭敏徐维祥

【Author】 Lei Shuo;Liu Xumin;Xu Weixiang;College of Information Engineering,Capital Normal University;College of Traffic and Transportation,Beijing Jiaotong University;

【机构】 首都师范大学信息工程学院北京交通大学交通运输学院

【摘要】 针对中文短文本词汇较少、噪声多、特征稀疏的特性,为了提高短文本分类精确度,提出一种基于维基百科词向量的特征扩展算法。利用维基百科语料集训练词向量,通过对文本关键词高相似度词集进行特征扩展,并将得到的文本用传统的分类器进行分类。实验结果表明,所提方法在短文本分类精确度上要优于其他的文本特征扩展算法。

【Abstract】 In order to improve the accuracy of short text classification,because of the characteristics of less,more noise and less characteristic of Chinese short text,a feature extension algorithm based on Wikipedia word vector was proposed to improve the accuracy of the text classification. By using the Wikipedia corpus training word vectors,the characteristics of the word sets of text keywords extended,and the text was classified by traditional classifier. The experimental results demonstrate that the proposed method is better than other text feature extension algorithms.

【基金】 国家自然科学基金项目(61672002);北京市长城学者项目(CIT&TCD20170322)
  • 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2018年08期
  • 【分类号】TP391.1
  • 【被引频次】44
  • 【下载频次】539
节点文献中: 

本文链接的文献网络图示:

本文的引文网络