节点文献

基于贝叶斯网络的XML文档检索

XML Documents Retrieval Based on Bayesian Network

【作者】 柴变芳

【导师】 徐建民;

【作者基本信息】 河北大学 , 计算机应用, 2006, 硕士

【摘要】 Web上的信息表示各种各样,为实现简单、方便、有效的Web查询,越来越多的信息采用XML数据模型来统一描述,XML将成为Web上数据描述和交换的标准,并且将会代替HTML而成为Web上驻留数据的主要格式。XML的出现为信息的处理提供了内容与结构两方面强有力的支持,对XML文档信息检索的研究也越来越受到重视。XML信息检索系统与传统的信息检索系统不同,主要体现在:建立索引时不仅需要建立倒排文本索引,还需要建立结构信息索引;查询处理时不仅需要处理查询中的关键字,还需要考虑文档的结构。 为满足大规模结构复杂的XML文档检索系统的需求,本文设计了一个检索策略:用户在简洁、一致的用户界面提交查询,系统对用户的查询进行优化处理,最后将文档按照满足查询的程度排序。主要取得4方面的成果:(1)索引机制的设计:根据XML文档的特点设计XML文档索引结构,并提出建立索引的算法:它对XML文档的各个元素内容和结构建立索引。(2)基于贝叶斯网络的XML文档查询选择模型:用户输入自然语言描述的查询后,系统根据文档集合的结构将其构造成多个结构化查询。对这些查询建立贝叶斯网络模型,通过模型推导出各个查询在当前文档集合下的概率公式。系统按照公式计算各个结构化查询的概率,一般概率最大的前n个最能代表用户的需求。(3)基于贝叶斯网络的XML文档排序模型:从(2)得到的每个查询都能返回大量的XML文档,对这些文档与查询建立贝叶斯网络模型,通过模型推导文档在查询下的概率公式。系统按照公式计算各个文档的概率,根据用户需求返回整个文档或者文档的某些部分的降序序列。(4)基于贝叶斯网络的XML文档检索原型系统:将论文中提出的新检索策略应用到原型系统,证明其可行性。并将文中检索策略与向量模型策略和信念网络模型策略进行比较,证明了它的准确性和有效性。 总之,本文的研究成果为建立高效的XML文档信息检索系统打下坚实的基础。

【Abstract】 The information on the Web is various. In order to realize a simple Web query efficiently and expediently, more and more information is described by XML data model consistently. And XML is quickly becoming the standard for data presentation and data exchange over the Internet, and is also specially designed for the web application instead of HTML. The emergence of XML sustains dealing with both the concepts and structures of the information forcefully. The researchers think much of XML document retrieval increasingly. XML information retrieval system differs greatly from traditional information retrieval system in the construction of both inverted text index and structural index, and the thought of both query keywords and the documents’ structures on dealing with user query.To manage large-scale XML documents with complicated structure, the article designs a retrieval strategy: user inputs their needs in a simple and unified interface. Then system optimizes the user’s query, and finally the results are ordered by the similarity between results and query. The paper mainly achieves the following contributions. (1) The indexing mechanism: The indexing structure is designed according to XML documents’ characteristics. Then the article provides an indexing algorithm which takes into account of the document’s structure and content. (2) The query selecting model of XML documents on Bayesian network: After user inputs the natural query, system constructs several structured queries according to the structure of the document collection. Then a Bayesian network model is built for all these queries, and the probability formula of each structured query given the document collection is inferred in the model. The system computes each structured query’s probability according to the formula, and the user selects several biggest queries that satisfy them. (3) The XML document ranking model on Bayesian network: A number of XML documents are returned from (2)’s query. Each document’s elements with the corresponding query are built to a Bayesian network model. According to the model the probability of each document on the

【关键词】 贝叶斯网络XML文档信息检索排序索引
【Key words】 Bayesian NetworkXML DocumentIRRankingIndexing
  • 【网络出版投稿人】 河北大学
  • 【网络出版年期】2006年 12期
  • 【分类号】TP391.3
  • 【被引频次】3
  • 【下载频次】189
节点文献中: