节点文献
基于贝叶斯的文本分类方法
Way of text classification based on Bayes
【摘要】 文本分类中的两个关键问题,算法和特征提取。贝叶斯算法是最有效的文本分类算法之一,但是属性间强独立性的假设在现实中并不成立,借鉴概率论中的多项式模型提出了一种改进型的贝叶斯方法;传统的特征抽取方法有词频法、互信息法、CHI统计、信息增益法等,然而上述方法对于词条的权重未作考虑,引进了权重的表征方式,给出了改进方法。由实验证明了通过以上方面的改进,文本分类的正确率得到了提高。
【Abstract】 Two important factors in text classification are discussed— algorithm and feature abstraction.The practical Bayesian algorithm has an assumption of strong independence of different properties and a modified way on polynomial is introduced.In Feature abstraction,different ways of abstracting features are discussed and a modified CHI based on word weight is introduced.At last the experiments show seen that correct rate of text classification is improved.
【关键词】 文本分类;
特征抽取;
贝叶斯;
多项式;
统计;
【Key words】 text classification; feature abstraction; Bayes; polynomial; statistic;
【Key words】 text classification; feature abstraction; Bayes; polynomial; statistic;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2006年24期
- 【分类号】TP18
- 【被引频次】46
- 【下载频次】844