节点文献
面向数据空间的分布式索引构建方法研究
Research on Distributed Index Construction Method for Data Space
【作者】 刘鹏;
【导师】 王念滨;
【作者基本信息】 哈尔滨工程大学 , 计算机科学与技术, 2021, 硕士
【摘要】 随着科学技术的进步,数据管理系统所面临的数据量越来越大。传统的关系型数据库逐渐无法满足快速增长的数据量。而且数据往往也不是由单一数据源组成的,而是分布在各个数据源。每个数据源之间的数据格式以及语义关系都不相同,用户需要消耗大量的时间和I/O资源对数据进行处理,无法快速从多源异构数据中获取有价值的信息。为了快速适应这种多源异构的数据环境,可以使用数据空间这一新的数据管理模式来解决当前的困难。以个人数据空间管理系统为例,用户不用再关注底层复杂多变的数据格式以及数据语义关系,直接高效的从数据中获取到有价值的信息。倒排索引在实际信息检索系统中被广泛应用,而如何利用倒排索引,快速从数据空间中多源异构的数据中,更高效获取有价值的数据,是当前索引架构的研究重点。本文通过对多种索引架构进行分析研究,提出了基于查询记录的分布式索引划分方法。对用户的历史查询记录进行挖掘,根据用户的查询偏好对高频检索词进行大小、负载可控的聚类,并根据每个节点处理能力不同,将高频词和缓存副本动态分配到各个处理机节点上,以及索引的查询策略和查询记录积累到一定程度后,倒排索引划分策略的动态更新调整策略,保证分布式索引系统各个处理机节点间的负载均衡,提升并行检索能力。将倒排索引划分到各个处理机节点后,随着每个处理机节点内的数据量累计增加,每个节点内token词数量过多,每个倒排列表的长度过长,降低查询性能。本文通过对传统倒排索引的水平分区和垂直分区的研究分析,提出了基于频繁模式挖掘的混合分区索引。本文提出了一种新的动态频繁模式树以及相应的创建和更新调整算法,来对传统频繁模式挖掘算法FP-group进行改进,提升了在数据更新时,发生频繁项集与非频繁项集转换而导致的结构更新的性能。在动态频繁模式树上挖掘出合适的token词进行垂直划分,然后在垂直划分的基础上进行水平划分,构建混合划分的倒排索引,提升索引的效率。
【Abstract】 With the development of science and technology,the amount of data faced by the data management system is increasing.Traditional relational databases are gradually unable to meet the rapidly increasing amount of data.And data is often not composed of a single data source,but distributed in various data sources.The data format and semantic relationship between each data source are different.Users need to consume a lot of time and I/O resources to process the data,and cannot quickly obtain valuable information from multi-source heterogeneous data.In order to quickly adapt to this multi-source heterogeneous data environment,data space,a new data management model,can be used to solve current difficulties.Taking the personal data space management system as an example,users no longer need to pay attention to the underlying complex and changeable data formats and data semantic relationships,and can directly and efficiently obtain valuable information from the data.Inverted indexes are widely used in actual information retrieval systems,and how to use inverted indexes to quickly obtain valuable data from multi-source heterogeneous data in the data space is the focus of current index architecture research.This paper analyzes and studies a variety of index architectures,and proposes a distributed index architecture method based on query records.Mining the user’s historical query records,clustering high-frequency search terms with a controllable size and load according to the user’s query preferences,and dynamically assigning high-frequency words and cache copies to each process according to the different processing capabilities of each node After the index query strategy and query records are accumulated to a certain extent,the dynamic update adjustment strategy of the inverted index partition strategy ensures the load balance among the processor nodes of the distributed index system and improves the parallel retrieval capability.After dividing the inverted index into each processor node,as the amount of data in each processor node cumulatively increases,there are too many token words in each node,and the length of each inverted list is too long,which reduces query performance.Based on the research and analysis of the horizontal partition and vertical partition of the traditional inverted index,this paper proposes a hybrid partition index based on frequent pattern mining.This paper proposes a new data structure dynamic frequent pattern tree and the corresponding creation and update adjustment algorithm to improve the traditional frequent pattern mining algorithm FPgroup,which improves the occurrence of frequent itemsets and infrequent items when data is updated.The performance of the structure update caused by the set conversion.The appropriate token words are excavated from the dynamic frequent pattern tree for vertical division,and then horizontal division is performed on the basis of the vertical division to construct an inverted index of mixed division to improve the efficiency of the index.
【Key words】 Data space; Inverted index; Load balancing; Partition index;
- 【网络出版投稿人】 哈尔滨工程大学 【网络出版年期】2022年 03期
- 【分类号】TP311.13
- 【被引频次】1
- 【下载频次】101