节点文献

流数据分析系统负载管理技术研究

Research on Load Management Technology for Data Stream Analysis System

【作者】 张奎

【导师】 王非;

【作者基本信息】 华中科技大学 , 通信与信息系统, 2013, 硕士

【摘要】 近年来,我国信息化进程取得极大进展。信息化的基础是数据的采集、存储、分析与利用。随着数据采集网络向更广更复杂的方向发展,同时数据采集的周期和采集精度不断提高,数据量呈现海量趋势。对于数据来讲,不仅具有值的属性,时间也是其重要的一个方面,数据的分析和利用与其时标特性密切相关,数据应用具有较高的时效性要求。多种采集环境、多种应用场景集成,数据形式、维度多样。总之,数据呈现海量、实时、多样的趋势。面对数据的新特点,传统数据库面临着诸多挑战和问题。首先,传统数据库存储所有的数据,在海量数据的应用场景下存储面临瓶颈;其次,传统数据库在数据查询存在大量的I/O操作,无法满足数据处理时效性的需求;再次,传统数据库无法适应数据分析的新需求。流数据分析系统是实时流数据采集、存储、分析的实时数据管理系统,在应对数据新特点带来的挑战方面有着诸多优势。实时流数据持续到达、速度快、规模大、不可预测,给流数据分析系统的负载管理带来了极大的挑战。流数据分析系统主要存在存储资源和计算资源两方面的性能瓶颈。目前的负载管理机制也是从这两个角度出发进行的。本文首先研究了负载管理的各种技术,核心是从计算资源的角度出发,以降低负载技术为理论基础,设计了一种负载管理算法。首先对流数据分析系统中所有的连续查询进行有向图建模,通过图中算子的选择率以及元组处理耗时计算查询网络的处理容量,进而判断过载时机,为了充分利用数据抖动的特性,减少降载的可能,提出了过载预测算法。基于连续查询的精确性描述,设计了降载的概率模型;为了实现服务质量的均衡,设计了基于降载优先级为核心的降载位置确定方法。通过仿真测试可以看出,在严重过载的情况下,算法降低了平均截止期措施率(Average Deadline Miss Ratio,ADMR),降低了查询结果的可用性损失(Utility Loss);同时仿真结果还显示,算法具有较好的自适应性、鲁棒性;在轻过载的情况下,本文的基于截止期的过载点预测算法很好的避免了实时降载方案,利用后续的处理空闲处理能力处理数据,保证查询的服务质量。

【Abstract】 In recent years, China made great process in the process of information. The basis ofthe process of information is the collection, storage, analysis and utilization of data. Withthe application develops widely and deeply, the widely spread devices of networks declaremore demand on data acquisition size, acquisition precision and acquisition speed. Theamount of data shows the trend of being massive. For the data itself, not only the value isimportant, but also its timestamp. Applications have a higher timeless requirement.Meanwhile, because of variety of acquisition environment and variety of applications, theform and dimension of data is variety.Traditional database has many problems to response to the new features of data.Firstly, the traditional database stores all the data, storage is a bottleneck cause of themassive of data. Secondly, when query data from traditional database, a large number ofIO operations are needed, the timeless of data process is hard to be meted. Thirdly, thetraditional database is unable to adapt to the new demands of stream data management andanalysis.Data stream analysis system(DSAS) is a real-time data acquisition, storage andanalysis system, there are many advantages to response to the challenges brought by thenew features of data. Real-time data stream arrives continuous, fast, unpredictably. Theload management of the system is a big challenge. Storage resources and computingresource are the main bottleneck of the system. Load management mechanism is carriedout based on these two aspects.This paper studies load management technologies, based on the computing resourceand load shedding theory, we design a load management algorithms. We design directedgraph model to describe continuous query in a DSAS firstly. Based on the selectivity andcost of different operators, algorithm estimates the capacity of the DSAS. With the changeof data stream rate, algorithm calculates the load of DSAS to decide whether to shed load.In order to utilize the shake of data stream, we design a method to expect the overloadpoint based on the deadline of queries. The key point of our algorithm is the method todecide where to shed load and the probability model of shedding load.In the experiments, our algorithm reduces Average Deadline Miss Ratio (ADMR) andUtility Loss when the system is severe overload. In the same time, our algorithm has goodadaptability and robust. However, when the system is in mild overload, the method toexpect the point of overload can use idle time followed-up to avoid implement of loadshedding and guarantee the accuracy of queries.

  • 【分类号】TP311.13
  • 【被引频次】1
  • 【下载频次】72
节点文献中: 

本文链接的文献网络图示:

本文的引文网络