节点文献
基于MapReduce的数据聚集运算算法
MapReduce-based data aggregation algorithms
【摘要】 为解决数据仓库中海量数据的处理效率问题,可以采用数据聚集预计算的方法,但是针对海量级别数据的聚集运算非常耗费计算资源,需要巨大的计算能力和存储能力,因此提出了一组基于MapReduce的面向海量数据的数据聚集运算算法,主要包括数据的选择、投影以及等值连接等,并在此基础上,实现了计数、求和和均值等聚集运算,形成了比较完整的面向海量数据的聚集运算算法。实验结果表明,该算法充分利用了集群系统的计算能力和存储能力,极大地提高了海量数据的聚集运算效率和基于聚集运算结果上的数据查询效率。
【Abstract】 To improve the computing efficiency of massive data in data warehouses,aggregation computing is one of the most typical data pre-processing methods.But it requires enormous computing power and storage capacity.So a set of MapReduce-based aggregation algorithms for massive data are proposed,mainly including data selection,projection and equivalent joint,etc.And the counting,summing,and averaging operations are implemented.They make a family of aggregation operation algorithms.Experiments show that the algorithms make full use of the cluster computing power and storage capacity,thus greatly improving the efficiency of the aggregation operations,and enhancing the query efficiency on massive data based on the aggregation results.
【Key words】 data warehouse; aggregation; MapReduce; on-line analytical processing;
- 【文献出处】 中国科技论文在线 ,Sciencepaper Online , 编辑部邮箱 ,2011年07期
- 【分类号】TP311.13
- 【被引频次】15
- 【下载频次】350