节点文献

基于句子级最大频繁单词集的Web文档聚类研究

Research on Web Document Clustering Based on Sentential Maximum Frequent Word Sets

【作者】 袁莉

【导师】 路松峰;

【作者基本信息】 华中科技大学 , 计算机应用技术, 2007, 硕士

【摘要】 Web文档聚类可以协助搜索引擎找出高质量的网页,是Web挖掘的一个重要研究方向。Web文档聚类技术的关键之一在于特征词或特征词组的选择。一篇文档的主题并不是与文档中的所有词相关,能体现文档主题的只有其中一部分,关键是要得到最能体现文档主题的特征项。现有的挖掘算法得到的频繁模式不仅维数高,而且不能很好的反映文档表达的语义信息。如何挖掘到理想的特征项,成为改进聚类算法的一个很重要的方面。为了解决这一问题,参考目前的数据挖掘领域的工作,给出了一个文档数据库模型,即将每一篇文档映射为一个数据库,文档中的每个句子看作文档中的一个交易,每一个词看作一个项目。然后利用关联规则挖掘算法来挖掘最能体现文档的特征单词集。相比较于传统的文档频繁特征项,句子级的频繁单词集包含了更多的局部信息。基于文档数据库模型,针对Web文档海量的特点,给出一种初步聚类和精确聚类相结合的两层聚类模型。先初步聚类后依据类间距离和类内链接强度阈值合并或拆分类,最后实现文档聚类。在此过程中,采用可变精度粗糙集模型计算文档中的每个频繁单词集对聚类的贡献,以此计算每个频繁单词集的权值。给出了基于容错粗糙集的聚类描述扩展。得到聚类结果后,为了进一步增强聚类的效果,对每个类别进行聚类描述。解决了由于同义词或者简写等语法现象造成的聚类描述不能精确匹配的问题,提高了聚类描述的有效性。容错粗糙集模型在处理模糊的、不确定关系方面有很大的优势。在信息检索领域中,特别是查询词扩展,文档与文档的关系,特征词与特征词的关系处理上得到充分的应用。

【Abstract】 Web document clustering could help search engineering to find out the web pages with high quality. It is an important research direction in web mining area. The one of the keys of the Web document clustering technique is the choice of the characteristic items. The theme of a document is not related of all word in the document, and the key is to find out the most feature items that reflect the themes of document. The dimension of the frequent pattern obtained from existing mining algorithm is high and not reflect the expression of semantic information well. How to mine the ideal characteristic items is becoming a very important facet of clustering algorithm.To resolve the problem, referring the current work in data mining field, present a document database model. Each document is mapped to a database, each sentence is regarded as a transaction and each word is regarded as an item. Then mining the characteristic word set who the most reflect the document by association rules mining algorithm. Compared with traditional frequent characteristic items, the frequent word set based sentence include more local information.In accordance with the feature of Web document with great quantity, present a two cluster model with initial clustering and precise clustering. After the initial clustering, merging or separating the classes based on the distance between two classes and links intensity threshold in a class, then achieving document clustering. In this process, compute the contribution of each frequent word set to clustering by using variable precision rough set and compute weight of each frequent word set.Present expansion of cluster description based on tolerant rough set. In order to intensify the effect of clustering, need to describe every class after acquiring the result of cluster. Because there are some syntax phenomenon such as synonyms or simplified versions in language. In order to express each class, need to expand the words. Tolerant rough set model have big advantage in processing fuzzy and uncertain relations. In the field of information retrieval, especially inquiries term expansion, the relation between documents and process the relation between the features with the full application.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络