节点文献
基于机器学习的文本分类方法研究
Research on Text Classification Method Based on Machine Learning
【作者】 于敏;
【作者基本信息】 江南大学 , 计算机技术(专业学位), 2021, 硕士
【摘要】 随着互联网时代的来临,每一刻都会产生海量数据,其中文本数据以传输效率高、便捷性高、普及范围广的优势存在于各个领域中,而如何对文本数据进行快速、准确的分类是当下的热门问题。本文以新闻文本为研究对象,对相关分类算法进行研究并改进,最终验证所提出的算法能够提高文本分类准确度。1.针对传统朴素贝叶斯文本分类算法中文本特征缺乏特征权重的问题,引入更侧重特征类别间分布的互信息,并将TF-IDF与互信息相结合,利用互信息关注特征词类别间关系的特点,补充TF-IDF的权重缺陷,并将改进后的方法所得到的权重融入朴素贝叶斯方法中,以减少传统方法中特征独立假设对分类的影响,提升分类器性能;2.针对传统的卷积神经网络文本分类模型没有对于学习到的文本特征进行区分,对于对文本分类结果意义大小不同的特征没有区别对待,所以引入注意力机制,即在卷积神经网络的全连接层前加入注意力层,将卷积池化层得到的文本特征进行注意力权重分配,使改进后的分类器更关注对于分类更有意义的特征,排除对于分类任务不重要的特征,实现分类效果提升的目的;3.针对中文文本篇幅较长,语法词语比较复杂的特点,本部分将通过在卷积神经网络中引入嵌套LSTM对模型进行改进。本文在局部特征提取的基础上,尝试对文本全局特征、上下文依赖关系进行提取,利用嵌套LSTM可以保存更长时间的记忆信息这一特点,引入嵌套LSTM以提取长时间的历史信息,更好地把握文本上下文语义,实现合理的新闻文本特征提取,提高分类准确率。最后,使用THUCNEWS新闻数据集、复旦新闻语料库和搜狗实验室新闻语料库文本数据进行实验验证。实验将改进后的贝叶斯分类模型与朴素贝叶斯分类模型做对比,将引入注意力机制与引入嵌套LSTM后的卷积神经网络分别与传统神经网络对比,根据准确率、精确度、召回率、F1值四个指标进行量化比较,结果表明本文所提出的算法模型能有效提高分类器性能。
【Abstract】 With the advent of the Internet era,massive amounts of data is generated at every moment.Text data exists in various fields due to its high transmission efficiency,high convenience,and wide range of popularization.How to quickly and accurately classify text data is a hot issue of the moment.This thesis takes news text as the research object,researches and improves related classification algorithms,and finally verifies that the proposed algorithm can improve the accuracy of text classification.1.Aiming at the problem of the same weight of every feature in the traditional naive Bayes algorithm,mutual information is used to focus more on the distribution of feature categories,which combines TF-IDF with mutual information,using mutual information to pay attention to the characteristics of the relationship between feature word categories,in order to supplement TF-IDF weight defects.And the weights obtained by the improved method are integrated into the naive Bayes method to reduce the influence of the feature independence assumption in the traditional method on classification and improve the performance of the classifier.2.For the problem of feature weights in the traditional convolutional neural network model,the attention mechanism is applied,that is,adding an attention layer before the fully connected layer of the convolutional neural network,and assigning feature weights to the text features obtained by the convolutional pooling layer,so that the improved classifier pays more attention for features that are more meaningful for classification,which can exclude features that are not important to the classification task to achieve the purpose of improving the classification effect.3.For the characteristics of the long length and the special word order of Chinese text,such as inversion,the convolutional neural network is further improved.Based on local feature extraction,this article attempts to extract global features and context dependencies of the text,using the feature that Nested LSTM can save memory information for a longer time.So nested LSTM is utilized to extract long-term historical information and grasp the text context and semantics better.Reasonable news text feature extraction is realized,and the classification accuracy is improved.Finally,THUCNEWS news data sets,Fudan news corpus and Sogou Lab news corpus text data are used for experimental verification.The experiment compares the improved Bayesian classification model with the naive Bayesian classification model,and compares the convolutional neural network after introducing the attention mechanism and the nested LSTM with the traditional neural network respectively.According to the four quantified and compared indicators which go as accuracy,precision,recall rate and F1 value,and the results show that the algorithm model proposed in this thesis can effectively improve the performance of the classifier.
【Key words】 Text Classification; Feature Weight; Naive Bayes; Convolutional Neural Network; Attention Mechanism; Nested LSTM;