节点文献

基于新的关键词提取方法的快速文本分类系统

Research on Fast Text Classifier Based on New Keywords Extraction Method

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 罗杰陈力夏德麟王凯

【Author】 LUO Jie,CHEN Li,XIA De-lin,WANG Kai (School of Eletronic Information,Wuhan University,Wuhan Hubei 430079,China)

【机构】 武汉大学电子信息学院武汉大学电子信息学院 湖北武汉430079湖北武汉430079

【摘要】 关键词的提取是进行计算机自动文本分类和其他文本数据挖掘应用的关键。系统从语言的词性角度考虑,对传统的最大匹配分词法进行了改进,提出一种基于动词、虚词和停用词三个较小词库的快速分词方法(FS),并利用TFIDF算法来筛选出关键词以完成将W eb文档进行快速有效分类的目的。实验表明,该方法在不影响分类准确率的情况下,分类的速度明显提高。

【Abstract】 Keyword extraction is the sticking point for Automatic Classification and Text Data Mining Application.Taking traits of nature language into consideration,this paper provides a new way called Fast Segmentation(FS) which is based on verb, virtual words and stop words to improve traditional segmentation technique.Then,we filter result of FS by TFIDF[3] Algorithm so that we can classify Web text fast and efficiently.The experiment has indicated that without reducing the correct rate of classification,the speed of processing has improved distinctly.

【基金】 国家自然科学基金资助项目(90204008)
  • 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2006年04期
  • 【分类号】TP391.1
  • 【被引频次】41
  • 【下载频次】831
节点文献中: 

本文链接的文献网络图示:

本文的引文网络