节点文献
分布式存储系统中磁盘故障检测机制
Disk failure detection mechanism in distributed storage systems
【摘要】 在大规模分布式存储系统中,经常会出现磁盘故障的情况,一方面需要尽快找出故障磁盘以降低数据丢失的风险,另一方面需要高准确率地找出故障磁盘以降低更换磁盘带来的时间成本和经济成本。文中针对以上需求,提出了一种基于磁盘空间随机取点的检测方法,通过将磁盘空间均分为N等份,然后在这些空间中随机读一个扇区,根据I/O状态以及I/O延迟时间来判断磁盘是否故障。实验表明,该方法能够在较短时间内以较高的准确率找出分布式存储系统中的故障磁盘,提高了分布式存储系统的可靠性。
【Abstract】 In large-scale distributed storage systems,there are lots of disk failures. On the one hand,it is necessary to find failed disks to reduce the risk of data loss. On the on other hand,it is necessary to accurately identify the fault disk to reduce the time costs and economic costs due to the replacement of the disk. For the above requirements,this paper proposes a method that accessing points randomly based on disk space. By dividing the disk space into N equivalents,then reading a sector in these equivalents.According to the I/O states and the I/O delay time to determine whether the disk is faulty. The experiments show that the method can find failed disks accurately in a short time in distributed storage systems,and improve the reliability of distributed storage systems.
【Key words】 distributed storage system; disk failure detection; random access point;
- 【文献出处】 信息技术 ,Information Technology , 编辑部邮箱 ,2018年05期
- 【分类号】TP333
- 【被引频次】3
- 【下载频次】115