节点文献

实时任务在集群计算中的自适应容错调度研究

The Study of Self-Adaptive Fault-Tolerant Scheduling for Real-Time Tasks on Cluster Computing

【作者】 王晓宇

【导师】 陆佩忠;

【作者基本信息】 复旦大学 , 计算机应用技术, 2010, 硕士

【摘要】 集群系统由于其良好的扩展性和可用性,逐渐成为当前并行计算的主要平台。随着实时应用范围的扩大,对计算机处理能力的要求不断提高,集群系统由于能够很好地处理计算密集型和数据密集型的应用,成为了解决这一问题的平台。在集群系统中,为了提升已有资源的资源利用率,调度算法的改进和优化是关键。由于实时任务有时间限制性的要求,因此为了防止系统出错而导致的任务失败,在实时任务调度算法中还必须考虑容错问题。本文在已有文献的研究基础上,研究异构集群中对于有服务需求的实时任务的容错调度算法,以提高集群系统的整体性能。已知的实时调度算法可以分为静态调度算法和动态调度算法两种。静态调度算法主要调度状态预先已知的任务,动态调度算法主要调度动态地到达的状态预先未知的任务。静态调度算法可以在编译时期实施完毕并且效率较高,主要用于调度周期性的硬实时任务;动态调度算法的灵活性更好,应用更加广泛。本文介绍了经典的静态调度算法及其扩展算法,还介绍了动态调度算法中任务的安全需求模型,系统的可靠性模型,以及系统的可获得性模型等研究内容。由于容错技术是实时任务调度算法的基本要求,本文介绍了Primary/Backup容错技术在调度算法中的应用。本文研究的调度算法是一种动态的调度算法。在基于异构集群平台上,为了处理有服务需求的实时任务,如具有安全需求或质量需求等服务需求的任务,本文采用Primary/Backup容错技术,综合考虑了任务的时间限制、任务的服务需求、系统整体性能等方面,提出了一种灵活的自适应容错调度算法SAOL。该算法在尽量满足系统对任务的调度成功率的基础上,根据系统的负载情况自适应地改变任务的服务级别。同时,为了减少容错技术带来的资源分配冗余,算法中还包含了PB和BB两种重叠技术,力求最大限度地提高资源利用率。本文经过模拟实验,将SAOL算法和已有文献中的算法做比较分析,实验结果表明SAOL算法具有更好的整体性能和调度灵活性。

【Abstract】 Owing to excellent extensibility and usability, heterogeneous cluster has gradually become the main platform of current parallel computing. With the growth of real-time applications, more and more computer resources of high computing speed and performance are required. Cluster computing can satisfy the need, because it is good at processing the applications of computing intensity and data intensity. In cluster system, it is key point to improve and optimize the scheduling algorithm for increasing the utilization of system resource. Because the time constraint of the real-time tasks, the scheduling algorithm must incorporate the fault-tolerance technique, for preventing from the task failure under the condition of system faults. Based on the theories of current literatures, this paper studies the fault-tolerance scheduling algorithm for scheduling tasks of service requirement in heterogeneous clusters, to improve the overall performance of the cluster system.The real-time scheduling algorithm of real-time tasks can be done either statically or dynamically. In the static scheduling algorithm, it schedules the tasks whose status is dertermined before the scheduling. In the dynamic scheduling algorithm, it schedules the dynamically arriving tasks whose status is unknown previously. The former can be applied at compiled time and is very efficient, and is mainly used to scheduling periodic tasks with hard deadlines. The latter is more flexible and popular in application. This paper introduces the classical static scheduling algorithm and the extended algorithms. It also introduces the task security requirement model, system reliability model, and system availability model of dynamic scheduling algorithms. Because the fault-tolerance is the essential requirement of real-time scheduling algorithm, this paper presents the application of Primary/Backup fault-tolerance technique in scheduling algorithms.In this paper, my study focuses on dynamic scheduling algorithms. On heterogeneous cluster platform, for dealing with the tasks of special service requirements, such as security requirement and service quality requirement, this paper proposes a flexible self-adaptive fault-tolerance scheduling algorithm, SAOL, using Primary/Backup fault-tolerance technique and considering the time constraint, service requirement, and system performance. The algorithm can adjust the task service level according to the burden of system, enhancing the flexibility and schedulability. Simultaneously, the algorithm incorporates the BB and PB overloading technique, for reducing the redundancy of resource allocation because of the Primary/Backup fault-tolerance technique, improving the resource utilization. At last, the algorithm is compared with an effective scheduling algorithm in the literature through simulation experiments. The experimental results show that the algorithm of SAOL has the superiority with higher performance quality and better flexibility.

  • 【网络出版投稿人】 复旦大学
  • 【网络出版年期】2011年 03期
节点文献中: