节点文献
融合语义特征的加权朴素贝叶斯分类算法
Weighted naive Bayes classification algorithm fusing semantic features
【摘要】 针对传统朴素贝叶斯算法属性独立性假设降低分类效果的问题,提出一种融合语义特征的加权朴素贝叶斯算法。在特征提取时引入Google距离衡量词语间语义相关性对节点权值进行重新计算;利用改进的NGD-TextRank算法提取数据集中关键特征,去除冗余属性进行降维;在分类过程中对不同特征项的影响程度进行划分,将特征项的权值融合到朴素贝叶斯公式中构造加权朴素贝叶斯分类算法。为验证算法性能,使用多类型数据集设计实验,与同类算法对比分析结果表明,该算法能够有效提取关键特征,经过加权处理后较好提高了文本分类的准确性。
【Abstract】 Aiming at the problem of reducing the classification effects of the traditional naive Bayes algorithm due to attribute independence hypothesis, a weighted naive Bayes algorithm for improving semantic features was proposed. The semantic distance between Google distance measurement words was introduced in the feature extraction to recalculate the node weights. The improved NGD-TextRank algorithm was used to extract the key features in the data set, and the redundant attributes were removed for dimensionality reduction. The degree of influence of different feature items was divided in the classification process, and the weights of the feature items were merged into the naive Bayes formula to construct a weighted naive Bayes classification algorithm. To verify the performance of the algorithm, the experiment was designed with multi-type datasets. The comparison with similar algorithms shows that the proposed algorithm can effectively extract key features and improve the accuracy of text classification after weighting.
【Key words】 Google distance; TextRank; Hellinger distance; weight optimization; naive Bayes;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2020年09期
- 【分类号】TP391.1
- 【被引频次】8
- 【下载频次】480