节点文献

一种基于维基百科知识库的中文文本分类方法研究

A Study of Chinese Text Classification Method Based on Wikipedia

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 苏小康何婷婷涂新辉何金卓

【Author】 SU Xiao-kang~(1,2) HE Ting-ting~(1,2) TU Xin-hui~(1,2) HE Jin-Zhuo~(1,2) 1(Department of Computer Science,Huazhong Normal University,Wuhan,430079;) 2(Monitor and Research Center for National Language Resource Network Multimedia Sub-branch Center,Wuhan, 430079)

【机构】 华中师范大学计算机科学与技术系国家语言资源监测与研究中心网络媒体分中心

【摘要】 传统的文本表示方法是基于词条的向量表示方法(Bag of Words or BOW),文本信息中的每一个词条都被表示成该向量中的一个维度。尽管这样的表示方法简单而且常用,但是却难免会有一些限制,因为文本之间存在着复杂的潜在的联系,而且这些潜在的联系很难用词条向量表示出来。因此在文本表示中插入一些背景信息用以提高文本分类模型的精确度是很必要的。该文通过搜集维基百科全书信息作为背景知识来扩充文本信息从而达到克服传统向量表示方法(BOW)的一些缺点,实验证明该方法可以提高文本分类的精确度。

【Abstract】 The traditional document representation is a word-based vector(Bag of Words,or BOW),where each term of the document is associated with a dimension of the vector.Although simple and commonly used,this representation has several limitations because of the complex relation between documents which BOW method can hardly deal with. To embed background information in order to enhance the accuracy of classification model is essential.In this paper,we proposed a approach of gathering the information of Wikipedia as background knowledge to enrich the information of documents,the method overcome the shortages of the traditional method(BOW).The model is proved that by use of Wikipedia to enrich the document information can improve the accuracy of document classification.

【基金】 国家自然科学基金(60773167);国家十一五科技支撑计划课题“网络文化安全预警技术研究”(2006BAK11B03);973国家重点基础研究发展计划(2007CB310804);教育部/国家外国专家局高等学校学科创新引智计划(B07042)
  • 【会议录名称】 中国计算机语言学研究前沿进展(2007-2009)
  • 【会议名称】第十届全国计算语言学学术会议
  • 【会议时间】2009-07-24
  • 【会议地点】中国山东烟台
  • 【分类号】TP391.1
  • 【主办单位】中国中文信息学会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络