节点文献
一种模仿人类的自动文本分类算法
An Automatic Algorithm of Text Categorization Imitating Human’s
【摘要】 <正> 1.引言 Internet上有着大量的且快速增长的文本,文本是信息和知识的宝贵资源。随着Internet的快速发展,不久的将来,人们所需要的大部分信息都可以在网上找到。Internet正在成为人类的信息宝库,但是随着网上信息的爆炸性增长,人们想从这个信息宝库中获得自己所需要的信息已经变得日益困难,因此,如何快速有效地获得有用的信息已成为人们十分关
【Abstract】 An algorithm of text classification is given that imitates human’s in this paper. On one hand, the algorithm enhances weight of theme when feature vector is processed, because of the assumption that the title of a document can project its content. On the other hand, a weight parameter to vector is designed to simulate human’s skimming and skipping behavior for calculating method of a document cluster center, and a weight of the feature that there are more positive examples than negative ones is enhanced . The experiment shows that the algorithm greatly improves the performance of a text classification system.
【Key words】 Text categorization; Corpus; Cluster center; Machine learning;
- 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2003年03期
- 【分类号】TP393.09
- 【被引频次】11
- 【下载频次】66