节点文献

基于概率的XML数据理论的研究

Research on XML Data Theory Based on Probability

【作者】 王建卫

【导师】 郝忠孝;

【作者基本信息】 哈尔滨理工大学 , 计算机应用技术, 2011, 博士

【摘要】 随着半结构化的概率数据的广泛应用,针对半结构化概率数据的理论研究是必要的。XML数据成为一种新的网络应用的数据形式,成为Internet中进行数据交换和表示事实上的标准的形势下,研究基于概率的XML数据理论具有较强的理论研究意义和应用价值。本文针对概率XML数据的管理问题,借鉴概率关系数据管理的思路和方法,对概率XML数据的管理理论涉及到的概率数据在XML数据中的表示方法、概率关系数据与概率XML之间的转换问题、建立概率XML代数操作集合、XQuery查询语言的概率操作扩充函数和元素节点的查询算法等几个方面进行了较系统的、较深入的研究。由于基于关系的概率数据是一种经典的概率数据形式,研究基于关系的和基于概率的XML数据的转换理论是有必要的。XML树和XML图是两种常用的XML数据模型,文中把基于概率的XML数据表示为概率XML数据树,提出了基于关系的和基于概率的XML数据的双向转换算法,该算法分为两个部分,一是模式转换,二是数据转换。在研究转换策略的基础上,提出了概率关系模式与概率XML模式PDTD的双向模式转换算法,并提出了概率关系数据转换为概率XML数据树和概率XML数据转换为概率关系数据两个数据转换的算法。在理论上对算法的正确性和完备性进行了证明,并通过与概率XML数据和概率关系数据的转换对比验证了该算法的正确性和完备性。设计概率XML数据的查询代数操作集合是实现概率XML数据库查询及查询优化的基本方法。将概率XML单元树作为概率XML数据代数的基本操作单位,其模式为概率XML模式树,设计了对遵循概率XML树模型的概率XML数据的集合的基本操作集合。给出了基于解析的路径表达式集合的各个基本操作的算法,在理论上对算法的正确性和完备性进行了证明,并通过实例验证了该算法的正确性。Xquery语言是XML数据的有效的查询语言之一,为了支持概率XML数据的查询,扩展Xquery的函数是一种简单的概率XML数据查询的实现方式。提出了扩展XML的查询语言Xquery函数的概率化的函数形式eXquery,按照扩展Xquery函数的功能分类的形式,设计了与路径表达式有关的函数、与节点有关的函数和与树类型有关的函数等。元素节点概率的查询是概率XML数据查询的主要内容之一,研究概率XML数据树的元素节点概率算法是必要的。在分析查询策略的基础上,提出了基于可能世界原理的查询算法和基于路径表达式集合的查询算法两大类算法,在理论上对该算法的正确性和完备性进行了证明,并通过实例验证了该算法的正确性,分析了算法的概率XML数据大小的适用性。

【Abstract】 With the comprehensive application of semi-structured probabilistic data, the research on the semi-structured probabilistic data management theory is essential.Because the XML data is a new data formal of network application and is the data interchange and the data representation standard on the Internet,the research on XML data management theory based on probability management theory has the theory research significance and the application impor-tance.Aiming at the management problem of probabilistic XML data, the research on the problems of the probabilistic data representation methods in the XML data, the research on data transformation problem between the probabilistic relation data and the probabilistic XML data, the design of the query algebra operation set, extending the Xquery function to query probabilistic XML data implementation, and the element node probability query problem are systemic and embedded.The probabilistic data based on relation is a classical probabilistic data formal, so the research on data transformation problem between the probabilistic relation data and the probabilistic XML data is essential. As XML data tree and XML data graph are two kinds of common XML data models, the XML data based on probability is denoted and the bidirectional transformation algorithm is represented between probabilistic data based on relation and XML.The algorithm is composed of two steps. The first step is the schema transformation procedure. And the second step is data transformation procedure. Based on the studied transformation strategy,the bidirectional probabilistic data schema transformation algorithms which convert probabilistic relation schema to probabilistic XML schema PDTD and convert probabilistic XML schema PDTD to probabilistic relation schema are proposed.The bidirectional probabilistic data transformation algorithms which convert probabilistic relation schema to probabilistic XML data and convert probabilistic XML data to probabilistic relation data are proposed. By theory analysis, it proves that the algorithms are right and self-contained. And it proves that they are right through the instances.The design of the query algebra operation set is the basic method of the probabilistic XML data query and query optimization, so the basic operation set of probabilistic XML data following the probabilistic XML data tree. In the basic operation set the basic operation unit is probabilistic XML unit tree whose schema is probabilistic XML unit schema tree. The algorithms of basic operation set based on the parsed path expression set are presented. By theory analysis, it proves that the algorithms are right and self-contained. And it proves that they are right through the instances.As Xquery is an effective XML query language, extending the Xquery function is an easy kind of query probabilistic XML data implementation method in order to support the probabilistic XML data query problem. The probabilistic function eXquery for the extended XML query language Xquery function set is presented. According to the function class of Xquery function set, the functions such as functions related with probabilistic XML tree types, path expression and node operation are designed.As the element node probability query problem is one of the main contents of probabilistic XML data query, so research on the element node probability of query algorithm in the probabilistic XML data tree is also necessary. Based on the analysis of query strategy, the two kinds of methods are presented. One is based on possible world model, the other is based on the parsed path expression set. During the course of the theory analysis, the algorithms are proved right and self-contained. It also proves that they are right through the instances and analyzed the applicability of probabilistic XML data tree size.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络