节点文献

地震数据处理系统的作业容错和迁移

Job Fault Tolerance And Migration of Seismic Data Processing System

【作者】 王建

【导师】 鲁才;

【作者基本信息】 电子科技大学 , 电子与通信工程(专业学位), 2014, 硕士

【摘要】 近年来,随着高性能计算持续高速发展,其结构组成越来越复杂,资源规模也越来越大,同时整个系统某个部件出现故障的概率也急剧增加。因此,利用计算机相关容错技术,为系统添加相应的容错功能,以便确保整个系统的高可用性,也显得越来越重要。现有的地震数据处理系统功能已经十分强大,支持复杂作业的处理执行,但是针对各种故障导致作业运行失败的情况,只是实现了最简单的容错处理,在实际使用中处理效率较低;同时在为作业选择运行节点时,采用了静态调度方式,在实际作业执行的过程中,各个节点资源仍然可能发生负载不均的情况,此时会降低系统的运行效率,加长了作业运行花费的时间。本文的目的是为现有地震数据处理系统,设计并实现一种支持作业自动容错和运行时作业动态迁移的子系统。通过深入研究计算机系统相关容错技术和进程迁移技术,在现有地震数据处理系统的基础上进行了功能添加和改进,主要工作包括:1、设计和实现了作业检查点自动容错子系统。通过深入理解现有检查点容错和恢复技术,针对集中式和分布式执行控制系统,设计和实现了用户级的检查点容错系统;针对具有道驱动特性的单道处理作业执行过程,设计和实现了应用级的检查点容错系统。2、设计和实现了作业运行时动态迁移子系统。在作业执行过程中,通过将作业进程从高负荷节点转移到空闲节点上继续执行,保障了整个集群数据处理系统的作业分布均衡,在静态调度的基础上进一步提高了集群系统的运行效率,保证了系统的负载均衡。3、实验测试。通过测试发现,作业检查点容错子系统中检查点设置和卷回恢复等操作花费的时间成本和存储开销都相对较小,不会影响作业的正常运行,有效地增加了数据处理软件的可用性;作业迁移子系统通过在不同节点间转移作业进程,进一步减少了作业执行花费的时间,保障了系统资源的有效利用。总之,本文的作业检查点系统能够有效解决由于故障导致的作业运行失败,极大的提高系统的容错处理效率,减少作业重复执行时间;同时通过作业迁移子系统的使用,有效的保证了集群的动态负载均衡,进一步提高了作业的执行效率。

【Abstract】 In recent years, with the rapid growth of high performance computing, its structure is more and more complex, and the scale of resource is becoming more and more big, at the same time, a component failure probability in the whole system is increased dramatically. Therefore, by using computer related fault-tolerant technology, adding the corresponding fault tolerance for the system, and ensuring the high availability of the whole system, also appear more and more important. The existed seismic data processing system is very powerful. It supports the processing of complex operation, but fails to run against all kinds of failure operation situation, and only the simplest fault tolerance is achieved, so processing efficiency is low in practice; Although using the static scheduling to select the node for job, in the process of the actual job execution, uneven load of each node resources may still happen, which lowers the system’s efficiency, and increases the job run time.The purpose of this article is designing and implementing a subsystems which is fault-tolerant and supports job dynamic migration for the existed seismic data processing system. Through in-depth study of related computer system fault tolerant technology and process migration technology, add and improve the function on the basis of the existed seismic data processing system, and the main work includes:First, this paper designs and implements the job checkpoint fault-tolerant subsystem automatically. Through the deep understanding of the existing checkpoint fault-tolerant and recovery technology, in view of the centralized and distributed execution control system, design and implement user level checkpoint fault-tolerant system; As for the driving characteristics of single channel processing job execution process, design and implement the application-level checkpoint fault-tolerant system.Second, this paper designs and implements the dynamic job migration subsystem. In the process of job execution, by moving job process from high load node to the idle node to continue execute, ensure the balanced of the whole cluster data processing system. on the basis of the static scheduling, improve the efficiency of cluster system, and ensure that the system of load balancing.Third, this paper does some experimental tests. By testing find that in the job checkpoint fault-tolerant subsystem checkpoint and rollback recovery operations cost little time and small storage, which does not affect the operation of normal job execution, and effectively increases the availability of data processing software; Job migration subsystem by transferring job process between the different nodes, further reduces the amount of time to finish the job, and ensures the effective utilization of system resources.In conclusion, this paper’s job checkpoint system can effectively solve the problem of job failure due to the breakdown of nodes, and greatly improve the efficiency of system fault-tolerant processing, reduce repeat time of job execution; At the same time, job migration subsystem can effectively guarantee the dynamic load balancing of cluster, and further improve the execution efficiency of jobs.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络