节点文献
基于高性能多流SSD的优化研究
Research of Optimization with High Performance Multi-streamed SSDs
【作者】 周波;
【导师】 诸葛晴凤;
【作者基本信息】 华东师范大学 , 计算机技术, 2021, 硕士
【摘要】 现如今,基于NAND闪存(NAND Flash)的固态硬盘(Solid state drives,SSDs)由于其高性能、低功耗的优势,被广泛的运用于数据中心,网络服务器,云计算等多个领域。与传统的存储设备不同,SSD存在着一系列独有的特性。数据更新在SSD内部以非就地更新的方式进行,并通过垃圾回收(Garbage collection,GC)操作回收无效的存储空间。但是频繁的GC操作将会影响SSD的使用性能,缩短SSD的使用寿命。为了降低GC操作对SSD的影响,多流(Multi-Stream)的概念被提出,它将不同生命周期,即冷热程度不同的数据存放在不同的流(Stream)中,并将相同的流的数据存储在SSD内同一闪存存储块中,从而保证同块闪存存储块中的数据几乎在同一时刻失效,进而降低GC操作的开销。而如何将不同生命周期的数据进行多流标定,是现如今多流SSD研究的一个热点。但现有多流标定方法都未考虑当数据以日志结构顺序写的方式更新时该如何进行有效的多流标定。另外,现有的工作都未考虑过如何基于多流技术设计SSD写缓存策略。带有多流信息的数据进入写缓存后,可能对彼此产生干扰,另外多流标定结果可能存在不准确的问题,这些都会影响多流SSD真实使用时的具体效果。本文针对高性能多流SSD设备的优化进行研究,包括数据日志顺序写方式更新场景下的多流标定方法的设计以及写缓存与多流技术相互作用的研究。本文的主要贡献如下:1)设计了一种基于F2FS文件系统(Flash Friendly File System)数据日志顺序写特征的多流标定方法。考虑F2FS文件系统的数据的日志顺序写更新方式以及数据的成块失效特征,基于文件系统的日志接入段进行多流标定。2)设计了一种多流感知的缓存隔离策略。基于写数据中带有的多流信息将归属于不同流的数据在写缓存中进行分区隔离,避免因不同流的数据在写缓存中互相干扰导致数据的冷热特性被破坏。该方法实现了基于多流技术来优化写缓存设计。3)设计了一种流内写缓存主动剔除策略,用于优化多流标定结果。考虑到部分数据的多流标定结果不准确的问题,主动地剔除写缓存中的部分数据。该策略能修正数据的冷热标定结果,使相同流的数据的生命周期特性趋于一致,提高多流标定结果的准确性。该方法实现了通过写缓存设计来优化多流技术。本文提出的多流标定方法实现了在文件系统层的进行多流标定,设计的两种写缓存策略实现了在写缓存层次优化现有多流标定方法的效果。本文通过在真实硬件平台以及模拟器上进行实验来验证所提出的方法的有效性。基于F2FS文件系统数据日志顺序写特征的多流标定方法的实验结果表明,在对大量小文件进行频繁小粒度随机更新的场景下,能使SSD的额外数据写入量降低800多倍,擦除次数降低4倍,使SSD的写放大下降到接近于1。多流感知的缓存隔离策略和流内写缓存主动剔除策略的实验结果表明,能在几乎不造成额外开销的前提下进一步地降低SSD的写放大。总的来说,本文所提的方法对于高性能多流SSD的优化工作具有重要的借鉴意义。
【Abstract】 Solid state drives(SSDs),which are constructed with NAND flash memory,have been widely adopted in data centers,web servers and cloud computing due to their high performance and low energy consumption.Different from traditional storage devices,SSDs have several limitations.NAND flash memory cannot do data update in place.To solve this issue,out-of-place update has been used by invaliding old data and writing new data to free space.However,once free spaces are consumed,a time consuming and lifetime impact process,called garbage collection(GC),is activated to reclaim the invalid data.GC will impact the performance and lifetime of SSDs.In order to optimize the GC process,multi-streamed SSDs have been proposed and widely developed.Its basic idea is to identify data with similar lifetime,we call it a stream,and write them to the same flash block group.The host can pass a stream ID along with write requests to the multi-streamed SSD,which convey hints on hotness of data.With this scheme,the overhead of GC process is significantly improved.One critical issue for the design of multi-streamed SSDs is to identify the data with similar hotness,which is called stream identification.However,there is no scheme works well when data are updated in an out-of-place manner.What’s more,none of previous works discussed how to efficiently utilize the stream information in the SSD controller,especially take write cache into consideration.When the write cache design is not aware of multi-streams,some problems may happen among streams and inside streams.First,data from different streams have different hotness,which may conflict with each other inside cache.Second,due to that it is a challenge to identify streams,the identification of streams may not be accurate.This paper focuses on the optimization of high-performance multi-streamed SSDs,including the design of stream identification scheme when data are updated in an outof-place manner and the study of the interaction between write cache and multi-stream technology.We summarize our contributions as follows.1)A stream identification scheme based on the append-only feature of F2 FS is proposed,which assigns different stream ID to the different log areas of F2 FS.This scheme can solve the problem that the logic space layout of data is not consistent with the physical space layout.2)A stream based write cache partitioning scheme is proposed to separate the management of data from different streams,which can solve the inter-stream conflict problem.This scheme can optimize the design of write cache with multi-stream information.3)An intra-stream based active cache evicting scheme is proposed to actively evict data to blocks with as many invalid pages as possible,which can solve the intra-stream inaccuracy problem.This scheme can optimize stream identification results through the design of write cache.We verify the effectiveness of the proposed schemes through experiments on real hardware platform and simulator.Experiment results show that when there are many small files with frequent granular random update,the proposed stream identification scheme can reduce the amount of extra writes to the SSD 800 times,erasing times 4times lower,make WAF dropped to close to 1.The proposed cache manage schemes are able to further reduce the write amplification of SSDs with negligible cost.In general,the schemes proposed in this paper has important reference significance for the optimization of high-performance multi-streamed SSDs.