节点文献
全局失效数据感知的键值存储系统优化
Optimization on Key Value Database with Global Invalid Data Awareness
【作者】 杨洁;
【导师】 童薇;
【作者基本信息】 华中科技大学 , 计算机技术, 2021, 硕士
【摘要】 基于日志结构合并树(Log-Structured Merge-tree,LSM-tree)的键值(Key-Value,KV)存储系统因其高性能和高扩展性得到了广泛的应用。由于LSM-tree键值存储采用异地更新,其不同层可能会存有同一个键(key)的不同版本键值对数据。LSM-tree键值存储通过对数据进行Compaction(归并排序)删除失效数据(除快照以外的旧版本键值对数据)。然而,每次Compaction只能根据参与的局部数据确定并删除部分失效数据,这可能导致失效数据长期滞留在系统中并多次参与Compaction,在影响系统读写性能的同时占用了系统的存储空间。对于数据更新分布均匀的应用而言,这些失效数据对系统产生的影响是不容忽视的。针对上述问题,设计并实现了全局失效数据感知的键值存储系统Gida DB(Global Invalid Data Awareness Database)。Gida DB根据LSM-tree内存只读组件中较新的数据,查找LSM-tree各层SSTable(Sorted String Table)中的旧数据,并将结果存储在旧数据信息表中。Gida DB并行执行查找旧数据与刷写内存只读组件,以减少旧数据查找带来的额外延迟。在Compaction时,Gida DB读取旧数据信息表用于检测参与Compaction的SSTable文件中的旧数据,并结合快照信息,识别并删除SSTable文件中的失效数据,从而减少失效数据多次参与Compaction引起的额外I/O,并减少失效数据带来的额外空间占用。测试结果表明,与Level DB相比,Gida DB的写放大降低了25.1%、写吞吐量提升了25.4%、空间占用减少了34.3%、读延迟减少了19.6%。在Level DB基础上同时部署Gida DB方案与最新写放大优化方案ALDC,与ALDC相比,Gida-ALDC的写放大减少10.7%、写吞吐量提升了8.9%、空间占用减少了22.1%、读延迟减少了13.0%,说明Gida DB方案能在一定程度上进一步提高ALDC方案的空间利用率和读写性能。
【Abstract】 Key-Value(KV)storage systems based on Log-Structured Merge-tree(LSM-tree)have been widely used because of its high write performance and scalability.Because LSM-tree Key-Value Store is updated in different places,different versions of key-value pair corresponding to the same key may be stored in different layers of LSM-tree.The LSM-tree Key-Value Store Engine removes invalid data(old version key-value pair data other than snapshot)by Compaction(merge sort).However,each Compaction can only determine and delete the invalid data based on the local data getting involved in Compaction,which may result in the invalid data staying in the system for a long time and getting involved in the following Compactions for many times,which degrades the system’s read and write performance and takes up excessive space of the storage.For applications that updates data uniformly,the impact of invalid data on the system cannot be ignored.To solve the problem above,Gida DB(Global Invalid Data Awareness Database),a KV storage system,is designed and implemented.Gida DB detects the old data in SSTables(Sorted String Table)of the LSM-tree(in layers L1~LN)using the newer data in the LSM-tree memory immutable component,and stores them in the old data information table.Gida DB creates the old-data-detection thread to do this,synchronized with the Compaction operation.During Compaction,Gida DB uses the old data information table to detect the old data in the SSTables getting involved in Compaction.Combined with the snapshot information,it identifies and deletes the invalid data in the SSTables,so as to reduce the extra I/O caused by the invalid data getting involved in Compaction for many times.And reduce the extra space used by the invalid data.The test results show that compared with Level DB,the write amplification of Gida DB is reduced by 25.1%,the write throughput is improved by 25.4%,the space consumption is reduced by 34.3%,and the read latency is reduced by 19.6%.On the Level DB basis,Gida DB is simultaneously deployed with the latest write amplification optimization scheme ALDC.Compared with ALDC,the write amplification of Gida-ALDC is reduced by 10.7%,the write throughput is improved by 8.9%,the space consumption is reduced by 22.1%,and the read latency is reduced by 13.0%.It turns out that Gida DB scheme can further improve the spatial utilization and read-write performance of ALDC scheme to some extent.
【Key words】 Key-Value Store; Log-Structured Merge-tree; Write Amplification; Invalid Data Awareness;
- 【网络出版投稿人】 华中科技大学 【网络出版年期】2022年 10期
- 【分类号】TP333
- 【下载频次】11