节点文献

数字图书馆中词频提取和自动文本分类方法的研究

【作者】 任美睿

【导师】 李建中;

【作者基本信息】 黑龙江大学 , 计算机软件与理论, 2002, 硕士

【摘要】 数字图书馆是一个新兴的、涉及到互连网、多媒体、数据仓库、数据挖掘、版权保护等诸多技术的计算机应用领域,应用和商业前景非常广阔。现在国内外对数字图书馆的研究刚刚起步。 我们在吸取前人经验的基础上,基于机群并行计算环境研制了一个并行数字图书馆系统,该系统除了具备现有数字图书馆的一些功能外,还可以根据用户的资源特点创建适合自己图书馆的元数据模式和分类体系模式。此外,该系统还提供了基于结构和内容的查询,这些功能是其它数字图书馆系统所做不到的。 本文设计并实现了数字图书馆中的词频提取和自动文本分类子系统,其中自动文本分类子系统克服了现有文本分类系统把文本类看作是互不相交的,处在一个平面层次上的弊端,依据数字图书馆中分类体系模式,实现了基于朴素贝叶斯原理的层次化自动文本分类。并提出了一个在特征提取阶段的有效的特征向量降维方法。在词频提取子系统中,本文根据中文词和英文词串的特点设计了一个高效的散列算法,这种散列方法能够较均匀地将文本中的词散列到散列表中,并快速定位到词的入口,有效提高了词频提取的效率。此外,本文还研究了基于向量空间模型的自动文本分类方法,提出了一个新的词权重计算方法,该方法有效提高了分类精度。

【Abstract】 Digital Library is a new computer application field that involves many technologies such as network, multimedia, data warehouse, data mining and copyright protection and so on, and research on it is at the beginning.A parallel digital library system based on parallel computing environment has been developed by our group. It has not only existing digital libraries’ general functions but also query function based on structure and content which isn’t realized in all other digital library systems. In addition, our system can establish adaptive digital libraries for our users with special needs.This paper designs and realizes the word frequency extract and automatic text categorization subsystem. Automatic text categorization subsystem can takes advantage of predefined class pattern’s hierarchical structure to construct hierarchical classifier, overcoming the shortcomings of other text categorization systems that consider classes flattening. In word frequency extract subsystem, the paper designs an efficient hash algorithm according to English words and Chinese words’ traits. The algorithm improves performance of the word frequency extract and statistics effectively. In addition, a text classification system based on Vector Space Model is studied and a new method for calculating word weight is proposed.

  • 【网络出版投稿人】 黑龙江大学
  • 【网络出版年期】2003年 02期
  • 【分类号】TP399
  • 【被引频次】7
  • 【下载频次】1004
节点文献中: 

本文链接的文献网络图示:

本文的引文网络