节点文献

基于概念格的Web文本聚类过程模型研究

Research on Clustering Process Model about the Text of the Web Based on Concept Lattices

【作者】 李海峰

【导师】 徐宝祥;

【作者基本信息】 吉林大学 , 情报学, 2010, 硕士

【摘要】 随着Internet时代的到来,Web信息呈现爆炸式增长,信息检索、文本挖掘这些情报学领域内的传统技术遇到了新的挑战,如何在海量级的Web文本中迅速找到符合自己需求的信息成为了一个新的课题。为解决这一问题,本文提出一种基于概念格的Web文本聚类过程模型,提供了一种全新的解决方法。本文以文本聚类现有的典型算法为切入点,深入分析了各个算法的优缺点,指出了现有文本聚类算法所存在的问题,接着引入了概念格这一新的数学工具,然后深入的了解了概念格的构造过程,并将这一构造过程引申到文本聚类过程中,使之成为了文本聚类的过程,从而在理论层面形成了一种全新的基于概念格的Web文本聚类过程模型。最后文本从应用角度验证了该模型,形成了2、3类文本聚类尤其是3类文本聚类结果与预想的结果相吻合,说明了该模型的可行性。本文的创新点在于给出了Web文本聚类的新方法,构造了一种基于概念格的Web文本聚类模型,在文本聚类方法上以及概念格应用上有所创新。文本的研究成果对情报学领域中的信息检索、文本挖掘提供了一种全新的解决思路,将数学领域内概念格方法进入了情报学领域进行研究,拓宽了研究思路,为情报学的相关研究起到了良好的促进作用。

【Abstract】 The growing of information technology especially the Internet has been challenging existing information phenomenon. And it has thrived World Wide Web’s information availability. The information explosiveness has been creating ways to find new concepts and mechanisms an effectively and efficiently to meet the information needs.To address the existing issues, this paper is proposed a web text clustering process model based on concept lattice.As an initial process, in-depth analysis of existing methods was carried to discover the advantages and disadvantages of each method. Followed by highlighted text clustering algorithms problems and then introduced a new concept lattice model which is a mathematical tool. And then the constructing of text clustering process and theoretical level formed presented. Moreover, the application point of view, two to three types text clustering is verified then demonstrated the feasibility of the model.Innovation of this paper is to give the new Web text clustering method to construct a concept lattice-based model of Web text clustering, text clustering methods and the application of concept lattices to innovate.the research results in the field of information retrieval, text mining solution offers new ideas, and also mathematics into the field of Information Science.It broaden the research ideas related to information science research played a good role in promoting.Having said that, the results of the research are presented as following:General introduction, background, significance and innovation task is presented in the chapter one. And chapter two is highlighted theories and concepts of given the theme like formal concept analysis, text clustering, clustering algorithm and web poly a class of process. In broader way, proposed model is presented in chapter three. Principles of Web text preprocessing, the text of feature representation and extraction methods whilst model background, extension table, building of lattice geometry, constructing ways of concept lattice of documents and properties of the cluster and also the formation of clustering results are respectively presented. The application process of the proposed model is discussed in chapter four. It has given overall surroundings of the entire application process. And finally, the findings, conclusions and further research are presented in chapter five.

  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2010年 09期
  • 【分类号】G350
  • 【被引频次】2
  • 【下载频次】256
节点文献中: