节点文献
数据流自适应查询处理技术
【作者】 周函;
【作者基本信息】 浙江大学 , 计算机应用技术, 2004, 硕士
【摘要】 随着网络技术的发展,越来越多数据正以数据流的形式存在于各种各样的网络系统中,同时以数据流处理为中心的应用也越来越多。因此,针对数据流的查询处理技术在近年来得到了学术界的广泛关注。由于数据流的处理与传统的静态数据的处理有着巨大的差异,传统的数据库管理系统(DBMS)已经不能适应对高速、时变的数据流处理的需求。因此,专门针对数据流处理的数据流管理系统(DSMS)也就应运而生了。 DSMS面临的重大挑战之一是长时间连续查询和多变的系统环境和数据所带来的对系统在适应性上的特殊需求。学术界在这方面还主要处在研究探索阶段。目前对数据流的自适应查询处理最为成熟和最有发展前景的成果是Eddy系统。本文以DSMS的适应性为研究重点,在继承Eddy系统在自适应方面的优越性的基础上,作出了两方面的新贡献。一是对用于在Eddy中自适应地处理多路Join操作的SteM机制作出了重大的改进,提出按需探测的中间结果算法,实现了在保持原有的SteM算法的适应性的基础上,对中间结果进行适当的保留,从而在Join的匹配率较高的情况下,减少了重复计算,提高了系统吞吐率,同时也对原有的路由策略也进行了针对SteM机制的有益的改进。另一方面,针对原有的Eddy系统的适应粒度过细的缺陷,本文实现了对Eddy系统的适应粒度进行自适应地控制,使得Eddy系统在数据和环境相对稳定的情况下,减少了路由决策的开销,提高了系统性能,同时并没有削弱在数据和环境变化频繁的情况下系统的适应性。本文详细地描述了上述两方面改进的具体算法和实现机制,对算法的性能进行了分析,给出了有说服力的实验结果,并指出了未来的研究方向。
【Abstract】 Recently, as the rapid evolution of the network technology, a new class of data-intensive applications has become widely recognized: applications in which the data is modeled best not as persistent relations but rather as transient data streams. Because data streams are continuous, unbounded, rapid and time-varying, traditional database management systems (DBMS) is not very suitable in processing this kind of data. So the data stream management system (DSMS) is being developed to focus on data stream processing.One of the biggest challenges in DSMS is the especial need for adaptivity, due to the long running continuous queries and time-varying feature of the data stream and the network environment. This is still a topic that has not been well studied in every detail. And to date, among the different prototypes in data stream query processing, the Eddy system is the most promising technology in adaptivity. The research of this paper is based on the Eddy, focusing on the adaptivity of query processing in DSMS. There are two significant innovations in this paper that make the Eddy more efficient and adapitive than the old one. First, we improved the SteM mechanism, which was used to deal with multi-joins in the Eddy. We proposed a method that can keep the intermediate results of the join operator, without reducing any adaptivity of the original SteM mechanism. And so more compute resources are saved to improve the whole throughput of the system. Second, as the granunarity is too fine in the original Eddy, especially in the circumstances that data and environments are relatively stable, we proposed a mechanism that can adaptively changing the granunarity of the adaptivity according to the varying rate of the data and the environment. So the system is more efficient in stable circumstances, without losing any benefit of the original adaptivity in unstable circumstances. All the details of the two contributions are described in this paper and promising results are given. We also indicated the future work at the end of this paper.
【Key words】 data stream; DSMS; adaptive query processing; multi-join; adaptive granunarity;
- 【网络出版投稿人】 浙江大学 【网络出版年期】2004年 02期
- 【分类号】TP393.09
- 【被引频次】1
- 【下载频次】204