节点文献
基于Spark的大数据统计中等值连接问题的优化
Optimization of the Equi-Join Problem Based on Big Data in Spark
【摘要】 伴随着互联网应用技术的飞速发展,导致传统的数据处理技术已经无法满足对大数据高效处理的要求。因此对现有的大数据的统计分析便急需相应的大数据技术的支持。为了解决实际Spark应用中的Join操作低效的问题,首先,提出一种高效的基于BloomFilter过滤再分区算法,通过该算法率先过滤掉绝大部分不符合条件的无效连接,然后针对过滤数据产生的倾斜问题进行再分区操作,以便能充分发挥各个工作节点的计算资源,达到在最大程序上优化Join过程的目的。
【Abstract】 With the rapid development of Internet application technology, leading to the traditional data processing technology has been unable to meet the requirements for the efficient processing of large data. So the existing large-scale statistical analysis of the big data will be in urgent need of the support of the corresponding large data technology. In order to solve the problem of the joining operation inefficient in practical Spark application, firstly, proposes an efficient algorithm of BloomFilter filter re-partitioning, uses the algorithm to filter out most of the invalid connections that do not meet the criteria. And then,re-partition the filtering data which is tilted, in order to make full use of the computing resources various nodes to achieve the maximum process to optimize the purpose of the join process.
【Key words】 Big Data; Spark; Equi-Join; BloomFilter; Shuffle;
- 【文献出处】 现代计算机(专业版) ,Modern Computer , 编辑部邮箱 ,2017年12期
- 【分类号】TP311.13
- 【被引频次】2
- 【下载频次】88