节点文献
文本特征加权方法TF·IDF的分析与改进
Analysis and improvement of feature weighting method TF·IDF in text categorization
【摘要】 TF·IDF作为一种简单、直观、处理速度快的文本特征加权方法,在文本分类中得到广泛应用。但是这种方法简单地认为文本频数少的单词就重要,文本频数多的单词就不重要,使它不可能很好的反映单词的有用程度,从而导致分类准确率下降。针对TF·IDF方法存在的问题,采用在特征发生的条件下类的后验概率分布来衡量特征对分类的有效性,提出了一种基于熵的特征加权方法TF·Ensu。实验结果表明,这种加权方法具有很好的分类性能。
【Abstract】 As a simple, direct-viewing, processing speed quick feature weighting method, TF · IDF method is widely used in document classification. But this method simply thought the words of low frequency are important, the words of high frequency are unimportant, which may decrease the precision of classification because of not reflecting the word useful degree. Entropy-based feature weighting method is presented, which solves the problems mentioned above. The method uses the category posterior probability-distribution to weigh feature to the classified validity. The experiments also show that it has good performance.
【Key words】 text categorization; feature selection; entropy; feature weighting; vector space model ..;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2008年11期
- 【分类号】TP391.1
- 【被引频次】37
- 【下载频次】684