节点文献

基于统计分布与集合论的文本分类方法

A Method of Text Classification Based on Statistical Technology and Set Theory

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 邓擘樊孝忠杨立公

【Author】 DENG Bo,FAN Xiao-zhong,YANG Li-gong(School of Computer Science and Technology,Beijing Institute of Technology,Beijing 100081,China)

【机构】 北京理工大学计算机科学技术学院北京理工大学计算机科学技术学院 北京100081北京100081

【摘要】 指出基于TfIdf的常用文本特征提取方法在文本分类问题中的缺陷,进而提出使用特征词的分布状态、词频和文本频三者相结合的方式提取文本特征的观点,给出了计算特征词权重的新方法,提出了新的文本分类方法.试验表明,该方法能够最大限度保留文本的特征,并且可有效避免向量空间模型中的维数灾难问题,能应用于大规模文本分类.

【Abstract】 Points out the limitations of general text feature extraction method based on TfIdf in problems of text classification,and presents the standpoint that combines the term distribution characteristic,term frequency and document frequency to extract the text feature,thus giving a new method to compute term’s weight,and a new way of text classification.Experiment showed that the method can keep the text’s feature to a maximum,and avoid the problem of dimensional disaster in VSM effectively,so it can be applied in problems of large scale text classification.

【基金】 教育部高等学校博士学科点专项科研基金资助课题(20050007023)
  • 【文献出处】 北京理工大学学报 ,Transactions of Beijing Institute of Technology , 编辑部邮箱 ,2006年07期
  • 【分类号】TP391.1
  • 【被引频次】11
  • 【下载频次】173
节点文献中: 

本文链接的文献网络图示:

本文的引文网络