节点文献

文本分类系统关键技术

Key Technologies of Document Classification System

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 谢科张辉陈鹏庞斌

【Author】 XIE Ke,ZHANG Hui,CHEN Peng,PANG Bin(National Key Laboratory of Software Development Environment,Beijing University of Aeronautics and Astronautics,Beijing 100083,China)

【机构】 北京航空航天大学软件开发环境国家重点实验室北京航空航天大学软件开发环境国家重点实验室 北京100083北京100083

【摘要】 从自然语言的角度考虑词性选择,同时从统计学角度考虑删除文档频率过低的特征词,从而避免产生维数灾难,通过考查类别本身特征和类别之间的关系来提取类别特征向量,采用传统夹角余弦公式考查文本与类别的相似度,实现一种过程简单,易于理解且分类效果不错的文本分类系统。

【Abstract】 From the view of natural language,POS selection is carried out;and from the view of statistics,the low frequency features are deleted.The relationship between the categories and their own features are examined to extract category features vector.Angle Cosine formula was used to examine the similarity between document and the category.A document classification system which is relatively simple,easily understood and works well has been presented.

【基金】 国家科技基础条件平台门户应用系统建设基金资助项目(2005DKA63901)
  • 【文献出处】 广西师范大学学报(自然科学版) ,Journal of Guangxi Normal University(Natural Science Edition) , 编辑部邮箱 ,2007年02期
  • 【分类号】TP391.1
  • 【被引频次】12
  • 【下载频次】252
节点文献中: 

本文链接的文献网络图示:

本文的引文网络