节点文献

多维数据存储及聚集优化策略的研究

Research on the Storage and Aggregation Optimizing Methods of Multidimensional Data

【作者】 熊东平

【导师】 蒋外文;

【作者基本信息】 中南大学 , 计算机应用技术, 2005, 硕士

【摘要】 数据仓库以传统的数据库为主要信息源,为联机分析处理(OLAP)、决策支持(DSS)和信息挖掘(DM)提供了一个集成的数据环境,高效地组织和管理数据是实现数据仓库技术的关键之一。本文从数据仓库的多维数据概念模型和OLAP实现两个方面对这个问题进行了深入的研究。 多维数据模型是数据仓库研究的核心问题。本文首先对现有的多维数据模型进行了总结,并分析了其优缺点,然后利用分类的方法提出了一个多维数据模型,并对数据立方体的相关概念进行了定义,该模型能够充分表达数据仓库的复杂数据结构,这为多维数据的存储及聚集优化策略奠定了理论基础。 多维数据的逻辑组织方式是OLAP实现的关键之一。本文对这个问题进行了深入的研究,总结了多维数据的两种组织方式—关系方式和数组方式,重点研究了在数组方式中多维数据的存储结构、多维数组的建立方法、稀疏数组的压缩方法、数组分块的原则和分块数组访问方法,并在以上的理论分析的基础上实现了一个数组方式存储的实例。 在数据仓库中高效计算多维聚集是提高OLAP性能的手段之一。本文总结了聚集计算的主要优化方法,对相关的概念进行了形式化定义,着重研究了数组聚集计算的优化策略,并给出了一个数组方式下的聚集算法——PartCube算法,该算法运用了最小父亲、阶段扫描以及缓存结果的优化策略,基于搜索格建立不完全方体的最小跨度树,当内存不足时,PartCube将数组进行划分并分别计算,计算完所有的划分后再把中间结果合并成完整的聚集结果。分析表明该算法能充分利用内存空间、减少I/O次数,具有较高的计算效率。 论文最后对研究工作进行了总结,并对进一步的研究工作进行了展望。

【Abstract】 The traditional databases are the main information sources of data warehouses; data warehouses provide an integrated data environment for Online Analytical Processing (OLAP), Decision Support System (DSS) and Data Mining (DM). Organizing and managing the data efficiently is one of the keys of implementing data warehouses. This thesis studies it deeply on the aspects of data warehouses’ concept model and OLAP implementation.Multidimensional data model is a basic aspect in the research field of data warehouses. After summarizing and analysing the existing data models, a data warehouses’ multidimensional data model is proposed using classification method and the corelative concepts of data cube are defined in this thesis, the model is powelful enough for modeling complex data structure of data warehouse, which establishes theoretical foundation for the storage and aggregation optimizing methods of multidimensional data.The logic organization mode of multidimensional data is one of the keys of OLAP implementation, this thesis summarizes the two organizing ways of multidimensional data - relational mode and array mode thoroughly, and places emphases on the researches of array mode, including the storage structure of multidimensional data, the construction methods of multidimensional arrays, the compressing methods of sparse arrays, the principles of dividing arrays into chunks and the access methods of chunk arrays, and also this thesis realizes a storage instance of array mode based on the above theoretical analyses.One means of improving the performance of OLAP is to compute multidimensional aggregations efficiently. This thesis summarizes the main optimizing methods of computing aggregations, on which the corelative concepts are formally defined, furthermore, this thesis emphasizes the research of optimizing methods of array mode and proposes an aggregation algorithm - PartCube algorithm, it makes use of optimizing methods including Small-parent, amortize-scans and Cache-results, and also it establishes the minimum spanning tree based on search lattice. If memory is insufficient, PartCube divides array into parts and computes each separately,after all parts have been accomplished, PartCube merges the intermediate results into integrated aggregations. The analysis shows that this algorithm can make the best use of memory and reduce I/O times, so it has high efficiency in computing aggregations.At the end of this thesis, the researches are summarized and the future work is presented.

  • 【网络出版投稿人】 中南大学
  • 【网络出版年期】2006年 05期
  • 【分类号】TP333
  • 【被引频次】4
  • 【下载频次】208
节点文献中: