节点文献
一种基于维基百科知识库的中文文本分类方法研究
A Study of Chinese Text Classification Method Based on Wikipedia
【Author】 SU Xiao-kang~(1,2) HE Ting-ting~(1,2) TU Xin-hui~(1,2) HE Jin-Zhuo~(1,2) 1(Department of Computer Science,Huazhong Normal University,Wuhan,430079;) 2(Monitor and Research Center for National Language Resource Network Multimedia Sub-branch Center,Wuhan, 430079)
【机构】 华中师范大学计算机科学与技术系; 国家语言资源监测与研究中心网络媒体分中心;
【摘要】 传统的文本表示方法是基于词条的向量表示方法(Bag of Words or BOW),文本信息中的每一个词条都被表示成该向量中的一个维度。尽管这样的表示方法简单而且常用,但是却难免会有一些限制,因为文本之间存在着复杂的潜在的联系,而且这些潜在的联系很难用词条向量表示出来。因此在文本表示中插入一些背景信息用以提高文本分类模型的精确度是很必要的。该文通过搜集维基百科全书信息作为背景知识来扩充文本信息从而达到克服传统向量表示方法(BOW)的一些缺点,实验证明该方法可以提高文本分类的精确度。
【Abstract】 The traditional document representation is a word-based vector(Bag of Words,or BOW),where each term of the document is associated with a dimension of the vector.Although simple and commonly used,this representation has several limitations because of the complex relation between documents which BOW method can hardly deal with. To embed background information in order to enhance the accuracy of classification model is essential.In this paper,we proposed a approach of gathering the information of Wikipedia as background knowledge to enrich the information of documents,the method overcome the shortages of the traditional method(BOW).The model is proved that by use of Wikipedia to enrich the document information can improve the accuracy of document classification.
- 【会议录名称】 中国计算机语言学研究前沿进展(2007-2009)
- 【会议名称】第十届全国计算语言学学术会议
- 【会议时间】2009-07-24
- 【会议地点】中国山东烟台
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会