节点文献
基于新的关键词提取方法的快速文本分类系统
Research on Fast Text Classifier Based on New Keywords Extraction Method
【摘要】 关键词的提取是进行计算机自动文本分类和其他文本数据挖掘应用的关键。系统从语言的词性角度考虑,对传统的最大匹配分词法进行了改进,提出一种基于动词、虚词和停用词三个较小词库的快速分词方法(FS),并利用TFIDF算法来筛选出关键词以完成将W eb文档进行快速有效分类的目的。实验表明,该方法在不影响分类准确率的情况下,分类的速度明显提高。
【Abstract】 Keyword extraction is the sticking point for Automatic Classification and Text Data Mining Application.Taking traits of nature language into consideration,this paper provides a new way called Fast Segmentation(FS) which is based on verb, virtual words and stop words to improve traditional segmentation technique.Then,we filter result of FS by TFIDF[3] Algorithm so that we can classify Web text fast and efficiently.The experiment has indicated that without reducing the correct rate of classification,the speed of processing has improved distinctly.
【Key words】 Computer Application; Nature Language Processing; Keyword Extraction; Web Text Classification;
- 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2006年04期
- 【分类号】TP391.1
- 【被引频次】41
- 【下载频次】831