节点文献

中文文本分类系统的研究与实现

The Research and Implementation of Chinese Text Categorization System

【作者】 甘立国;

【导师】 董小国;

【作者基本信息】 北京化工大学 , 计算机应用技术, 2006, 硕士

【摘要】 随着信息技术的迅速发展,特别是Internet的普及,网页数量呈海量增长。由于网页中的内容大部分是文本信息,因此如何根据网页中的文本信息自动分类成为目前研究的重要课题。文本自动分类是信息检索中的一个重要环节,它是指在给定的分类体系下,根据文本的内容自动判定文本类别的过程,以便于信息的检索。本文首先介绍了文本自动分类在国内外的研究现状;其次对文本自动分类所涉及的关键技术,包括信息检索模型、中文分词方法、特征抽取、特征项权重方法以及关键的分类算法,分别进行了研究和探索;再次在特征项权重方面,我们分析了传统特征项权重方法的缺点,提出使用句子的重要度对特征项的权重进行加权,实验证明这种方法能有效地反映文本的内容;接下来介绍了基于向量空间模型的中文文本分类系统的总体框架,系统流程和功能模块;最后对分类系统中实现的各种特征抽取算法、权重算法和分类算法分别进行了实验对比。

【Abstract】 With the development of Information technology and the prevalence of Internet, the amount of web page increase explosively. Because the content of web page is mostly text, how to categorize web page automatically by its text information became an important research subject. Text categorization, the automated assigning of natural language texts to predefined categories based on their contents, is an important part of Information retrieval. This paper firstly introduce the research status of text categorization, secondly we study and discuss the key technique of text categorization, including Information retrieval model, Chinese word segment, Feature Selection, Feature Weight and Classify Methods. Considering the disadvantage of tradition Feature Weight, we use sentence’s importance to compute feature’s weight and experiment prove that this method is good for Categorization. Thirdly, we introduce the frame, system flaw and function module of Chinese text categorization system based on vector space model. Finally, we list the result of experiment on feature selection, feature weight and classify

  • 【分类号】TP391.1
  • 【被引频次】38
  • 【下载频次】429
节点文献中: