节点文献

一个基于语义信息提取的互联网情报挖掘系统的设计与实现

A Design and Implement of Internet Intelligence Mining System about Semantic-based Information Extraction

【作者】 黄朝晖

【导师】 姜晓红; 陈华钧;

【作者基本信息】 浙江大学 , 计算机应用技术, 2010, 硕士

【摘要】 随着Internet的高速发展,Web已经成为世界上规模最大的公共数据源。人们可以从Web获取信息,可以通过Web与其他人交流,可以在Web上共享自己的信息。然而由于Web数据规模如此庞大,如何从中快速准确的检索到用户所需要的信息是一个急迫需要解决的问题。针对这一问题,在信息检索领域中的Web数据挖掘便应运而生,并且伴随着Web的发展而备受关注。Web数据挖掘它建立在信息检索、数据挖掘以及知识管理等技术的基础上,通过对大量的Web文档进行分析来获得隐含的知识和模式,从而帮助人们更好的进行信息检索和决策制定。本文分析了Web数据挖掘的研究内容和研究状况,设计并实现了一个基于语义信息提取的互联网情报挖掘系统,具体的内容包括:1.实现并分析了Web页面提取、网页正文提取、自然语言处理以及关键字信息抽取等子系统模块;2.提出并实现了语义关系图的构建模型,该模型用图的形式表示非结构化的文本数据中的语义关系;3.实现了一种频繁子图挖掘算法,该算法不同于单纯的深度遍历和广度遍历算法,存效率上优越于前两者;本文将该算法应用于挖掘潜在的频繁语义子图,得到具有一定客观性的语义关系图;4.提出并实现了一种基于Linked Data的RDF链搜索算法,用Linked Data解析频繁子图,从而获得具有标注关系的语义关系图。

【Abstract】 With development of Internet, web has become the biggest open data resource in the world. People can achieve information from web, connect others by web and share their resource on web and so on. But the web resource database are such large that how to get the information satisfied with user’s demand quickly and exactly is an urgent problem. To solve this problem, a new technology named "Web Data Mining" was introduced in information retrieval domain, and it was paid much attention by investigators with the development of web. Web data mining is built on the base of information retrieval, data mining and knowledge management, and achieve implied knowledge and pattern by analyzing large number of web documents, so that it can improve information retrieval and decision making.This paper analyze the recent investigated content and progressivity in the domain of wed data mining, then design and realize a web information mining on semantic-based information extraction. The concrete content includes as follow:1. Implement and analyze some subsystem modules, such as web pages crawling, main contents extraction from web pages, natural language processing and key words extraction.2. Put forward and implement a semantic relational graph construction model, which employ graph to express semantic relationship in non-structured text data.3. Realize a frequent subgraph mining algorithm, which is different with DFS and BFS algorithms, and is more efficient than them. This paper employ the algorithm for mining frequent semantic subgraph, then get some objective results.4. Inventing a RDF searching algorithm on Linked Data, which is used to explain frequent subgraph, then get the results with relationship label.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2011年 04期
节点文献中: