节点文献
分布异构集群下的MapReduce作业调度
Mapreduce Job Scheduling for Heterogeneous Geo-distributed Clusters
【作者】 王佳;
【导师】 李小平;
【作者基本信息】 东南大学 , 计算机应用技术, 2019, 博士
【摘要】 分布异构集群下的MapReduce作业(独立MapReduce作业或MapReduce工作流)调度的主要问题是任务与资源间的合理匹配。实际云环境中资源的有限性、MapReduce作业处理数据的分布性以及不同类型MapReduce作业资源请求量的异构性,为MapReduce作业或MapReduce工作流调度过程中满足截止期、数据本地化、资源利用率等带来极大挑战。本文围绕MapReduce作业调度,分别结合最大完工时间、能效优化和收益代价展开研究调查,设计了基于分布异构集群的时间感知MapReduce作业调度方法、基于分布异构集群的能耗感知MapReduce作业调度方法和基于分布异构机器的收益感知MapReduce工作流调度方法。主要工作如下:1、基于分布异构集群的时间感知MapReduce作业调度:考虑数据本地化、作业截止期和自适应心跳等约束以最小化多个独立MapReduce作业的最大完工时间为优化目标。将map任务分配到分布式集群处理以减少数据传输时间;reduce任务处理考虑中间数据传输时间和任务执行时间,选择具有最早完成时间的集群,减少MapReduce作业的完成时间;根据map任务与reduce任务最大处理数据量的比例,将MapReduce作业截止期划分为map阶段截止期和reduce阶段截止期;将该问题建模为指派问题。依据作业各阶段的任务处理时间计算自适应心跳间隔,MapReduce作业在每个心跳周期内按照截止期排序,利用匈牙利算法将任务分配至合适的资源槽以减少作业完成时间。实验结果表明,相比于已有算法,所提算法可得到更好性能。2、基于分布异构集群的能耗感知MapReduce作业调度:考虑作业截止期、数据本地化和资源利用率等因素以最小化分布异构集群的能量消耗为优化目标。对该问题建模并提出一个动态MapReduce作业调度框架。MapReduce作业根据作业截止期、作业可分配的资源槽数目和作业预估执行时间排序;不同任务从相应的机架层本地机器、集群层本地机器和远程机器中选择最有价值的资源槽分配以改善数据本地化;计算集群中可用资源槽的更新除查找可用资源槽外,还要根据节点当前的CPU、内存和带宽利用率采用模糊逻辑动态改变节点中资源槽的数目以提高资源利用率。实验结果表明,提出的启发式算法所消耗的计算集群能量要少于已有算法。3、基于分布异构机器的收益感知MapReduce工作流调度:考虑工作流截止期和数据本地化等约束以最大化资源管理者的收益代价为优化目标。对该问题提出改进的工作流调度架构、给出数学模型和一个工作流调度框架。依据已有Chain Map/Chain Reduce,MapReduce工作流利用动态规划转换以合理减少数据传输时间;工作流截止期根据MapReduce作业的预估执行时间、作业浮动区间[1]和作业层级划分为MapReduce作业截止期;按照工作流、MapReduce作业和任务的调度序列给出4种不同的任务列表构造方法;MapReduce工作流调度引入复本策略以改善数据本地化;将任务分配至具有最小完成时间的机器来增加资源管理者的收益。实验结果表明,相比于已有策略,提出的工作流调度算法使得资源管理者可获取更多的收益。
【Abstract】 The main problem of MapReduce job/workflow scheduling is to assign tasks to server-s reasonably in heterogeneous geo-distributed clusters.Because of the heterogeneous,geo-distributed and limit servers in cloud center,the random distribution of data and the hetero-geneous resource requirements of MapReduce job/workflow,many challenges are brought to MapReduce job/workflow scheduling with deadlines,data locality and resource utilization.In this paper,the MapReduce job/workflow scheduling in heterogeneous geo-distributed clusters with minimizing makespan,with minimizing energy consumption and with maximizing ben-efit cost are considered respectively.The main contributions of this paper are summarized as follows:1、A time-aware MapReduce job scheduling in heterogeneous geo-distributed clusters is considered to minimize makespan with deadlines,data locality and adaptive heartbeat interval.Map tasks of jobs are processed in parallel in different clusters to decrease data transmission times.Reduce tasks of jobs are processed in one cluster with the minimal estimated completion times of jobs according to shuffle times of intermediate data and processing times of reduce tasks.In terms of the maximum data volume of tasks,deadline of MapReduce jobs are divided into deadline of map tasks and that of reduce tasks.The MapReduce job scheduling is formulated as an Assignment Problem,in which adaptive heartbeats are calculated by processing times of tasks.In each heartbeat,jobs are sequenced in terms of the divided deadlines and tasks are scheduled by the Hungarian algorithm to decrease completion times of jobs.Experimental results show that the proposed algorithms outperform the existing works.2、An energy-aware MapReduce job scheduling in heterogeneous geo-distributed clusters is considered to minimize energy consumption with deadlines and data locality.The MapReduce job scheduling is modeled and a dynamic MapReduce job scheduling framework is proposed.Jobs are sequenced according to deadline constraints,allocated number of job slots and possible processing times of jobs.Tasks are scheduled to promising slots from their rack-local servers,cluster-local servers and remote servers in order to improve data locality.An update of available slots in clusters is proposed not only to find available slots but also to improve server resource utilization using fuzzy logic with the available number of slots according to current CPU,memory and bandwidth utilization.Experimental results show that the proposed heuristic results in lower energy consumption than adopted algorithms from literatures with a variable total number of slots.3、A benefit-aware MapReduce workflow scheduling in heterogeneous cloud center is considered to maximize benefit cost of resource managers with deadlines and data locality.The MapReduce workflow scheduling is modeled and a modified architecture of workflow scheduling is designed.Meanwhile a workflow scheduling framework consisting of workflow conversion,deadline division,task list construction and task scheduling is proposed.A number of MapReduce workflows are converted by Dynamic Programming according to Chain-Map/Chain Reduce in order to decrease transmission times among jobs reasonably.In terms of execution time,float time[1]and job level,deadlines of these converted workflows are divided into subdeadlines of jobs in workflows.Four different task list constructions are designed according to the sequences of workflows,MapReduce jobs and tasks.In order to improve data locality,replica strategy is adapted in MapReduce workflow and tasks are scheduled to servers with the earliest finish time.Experimental results show that the proposed heuristic results in more total benefit than other adopted algorithms.
【Key words】 MapReduce job; Heterogeneous environment; Deadline; Data locality; Resource utilization;