节点文献
基于改进TF-IDF特征提取的文本分类模型研究
Research of Text Classification Model Based on the Improved TF-IDF Feature Extraction
【摘要】 【目的/意义】特征提取会很大程度地影响分类效果,而传统TF-IDF特征提取方法缺乏对特征词上下文环境和对特征词在类之间分布状况的考虑。【方法/过程】本文提出一种改进TF-IDF特征提取的方法:(1)基于文本网络和改进Page Rank算法计算节点重要程度值,解决传统TF-IDF忽略文本结构信息的问题;(2)增加特征值IDF值的方差来衡量特征词w在不同类别文本集中程度的分布情况,解决传统TF-IDF忽略特征词在类之间分布状况的不足。【结果/结论】基于该改进方法构建了文本分类模型,对3D打印数据进行分类实验。对比算法改进前后的分类效果,验证了该方法能够有效提高文本特征词提取的准确度。
【Abstract】 【Purpose/significance】Feature extraction plays an important role in text classification, while traditional TF-IDF method lacks consideration of the context of feature words and its distribution between the classes.【Method/process】The study proposes an improved TF-IDF feature extraction methods: 1) in order to solve that the traditional TF-IDF ig-nores the text structure information, the paper computes node importance value based on text network and improved Page R-ank algorithm; 2) in order to solve that the traditional TF-IDF overlooks feature words distribution between classes, the pa-per increases the variance of IDF values represent the distribution of text focused concentration of different types of w.【Re-sult/conclusion】Based on the improved method to construct a text classification model, and take 3D printing as a classifica-tion case. Comparing the classification results before and after the improved algorithm process, the improved TF – IDFmethod is verified to extract text feature words accurately and effectively.
【Key words】 feature extraction; TF-IDF; text classification; text network; PageRank;
- 【文献出处】 情报科学 ,Information Science , 编辑部邮箱 ,2017年05期
- 【分类号】TP391.1
- 【被引频次】104
- 【下载频次】1478