节点文献

基于语义扩展的短文本分类研究

Research on Short Text Classification Based on Semantic Extension

【作者】 李珍

【导师】 刘怀亮;

【作者基本信息】 西安电子科技大学 , 情报学, 2019, 硕士

【摘要】 互联网的迅速发展加剧了信息时代的进步,短文本作为一种简单高效的表达方式广泛存在于各种社交网站中,如微博、新闻标题、商品评论、论坛、朋友圈等,想要从这些海量的文本资源中挖掘出有用的信息变得愈加困难。由于短文本具有稀疏性、即时性、海量性、不规则性等特点,传统的分类方法仍然存在文本语义信息提取不足和严重的数据稀疏问题。目前,引入外部知识库来扩展短文本的语义信息是较为热点的研究方向,如何才能获得文本中多层语义表达,并且消除短文本中不相关术语的影响,成为当前短文本分类研究的一个重要课题。针对上述问题并参考已有的研究成果,本文引入语义特征扩展的思想,将Probase语义网络作为外部知识库,通过词语概念化和增加语义共现词的方式对短文本进行扩充,使其能够更好地表达短文本中隐含的信息,达到消歧的效果。然后再结合Word2vec模型训练语义信息词向量,很好地解决了文本表示所面临的数据稀疏性和词语之间语义不足的问题,在传统分类模型的基础上,提出了一种基于语义扩展的短文本分类方法。本文首先仔细分析了短文本独有的特点和传统短文本分类技术,指出了传统短文本分类模型中存在的缺陷,确定了Probase知识库相较于其他知识库在扩展短文本语义信息上的优势;其次,推断出短文本中每一个词语符合该语境的概念词和共现词,然后作为词语的语义信息添加到文本中,同时根据上下文语境选取最具代表性的概念进行匹配,并删除模糊术语。结合Probase语义网络和Word2vec词向量对文本进行特征向量表示,该方法不仅能够丰富短文本语义信息,而且还能准确地表现出词语之间的相互联系以及上下文结构表达;再次,针对传统分类模型,从短文本预处理、文本表示等步骤进行优化,概念化的短文本采用基于Word2vec模型的短文本分类方法解决传统分类模型中存在的文本特征向量维度过高和稀疏性的问题,获得高质量的语义特征词向量表示;最后,通过比较目前已有的分类方法,选择LIBSVM算法进行短文本分类,将本文提出的基于语义扩展的短文本分类方法与传统的分类方法进行对比。实验结果表明,本文所提出的方法可以取得更好的分类效果。

【Abstract】 The rapid development of the Internet has intensified the progress of the information age.Short texts exist as a simple and efficient expression in various social networking sites,such as weibo,news headlines,product reviews,BBS,circle of We Chat friends etc,It is becoming more and more difficult to extract useful information from these massive text resources..Due to the sparseness,immediacy,mass and irregularity of short texts,The traditional classification method still has insufficient text semantic information extraction and serious data sparseness.At present,the introduction of external knowledge base to extend the semantic information of short text is a hot research direction.How to obtain multi-layer semantic expression in text and eliminate the influence of irrelevant terms in short text has become an important research of short text classification.question.In view of the above problems and referring to the existing research results,this thesis introduces the idea of semantic feature expansion,and uses Probase semantic network as external knowledge base to expand short text by word conceptualization and increase semantic co-occurrence words to make it better.Express the information implied in the short text to achieve the effect of disambiguation.Then the Word2 vec model is used to train the semantic information word vector,which solves the problem of data sparsity and semantic deficiency between words.Based on the traditional classification model,a short text classification method based on semantic extension is proposed.This theisis firstly analyzes the unique characteristics of short texts and traditional short text classification techniques,points out the shortcomings of traditional short text classification models,and determines the advantages of Probase knowledge base in extending short text semantic information compared with other knowledge bases.Secondly,it is concluded that each word in the short text conforms to the conceptual word and co-occurrence word of the context,and then is added to the text as the semantic information of the word,at the same time,according to the concept of context to select the most representative matching,and remove the fuzzy terms,Combining Probase semantic network and Word2 vec word vector to represent the eigenvectors of text,this method can not only enrich the semantic information of short text,but also accurately represent the interrelationship between words and the expression of context structure;Thirdly,the traditional classification model is optimized from text preprocessing,text representation and other steps.The conceptualized short text classification method based on Word2 vec model is adopted to solve the problem of high dimension and sparsity of text feature vectors in the traditional classification model,so as to obtain high-quality vector representation of semantic feature words.Finally,by comparing the existing classification methods,the LIBSVM algorithm is selected for short text classification,and the short text classification method based on semantic extension proposed in this thesis is compared with the traditional classification method.The experimental results show that the proposed method can achieve better classification results.

【关键词】 短文本Probase词向量特征扩展
【Key words】 short textProbaseWord2vecFeature extension
节点文献中: