节点文献
数据空间数据采集存储与可用性管控技术研究及系统实现
Research and System Implementation of Data Space Data Collection,Storage and Availability Control Technology
【作者】 蒲长春;
【导师】 吴奇石;
【作者基本信息】 西南交通大学 , 计算机技术(专业学位), 2023, 硕士
【摘要】 现代社会信息技术高速发展,使得网络与软件无所不在,人与人的交流带来了大量的数据传输与流通,数据已然成为当今社会不可或缺的“血液”。随着数据量的激增,汽车产业链上的数据处理技术迎来了巨大挑战,数据质量保障能力出现了严重问题,由此引发了关于汽车行业大型数据处理平台在数据采集、存储、可用性等方面的一系列严峻考验。为此,本文面向汽车产业链数据空间,研究了数据空间数据采集存储与可用性管控技术,旨在加强数据空间的大数据管理与分析能力,提高数据空间智能服务的质量。本文首先分析了数据空间数据采集存储与可用性问题及相应的管控需求,总结了现有技术方案的不足,在此基础上提出了数据空间数据采集存储与可用性管控解决思路。接着,基于该解决思路,论文设计了数据采集存储与可用性管控系统的总体架构及各个功能流程,同时设计了系统数据库的概念模型、逻辑模型和数据表结构。基于Hadoop生态的大数据处理技术,本文提出了适配数据空间的分布式数据采集存储基础架构,并进一步构建了数据空间数据可用性管控模型。对于基础架构,基于Flink+Hudi增强了数据空间的实时与离线采集能力,基于HDFS+Hive+Hudi增强了数据空间的大数据存储能力和多计算引擎支撑能力;对于可用性管控模型,从动态流转数据的角度提出了流转数据可用性量化指标并设计了量化指标的计算方法,从静态存储数据的角度提出了可用性提升策略。其中,对于可用性提升策略,设计了基于约束条件、LOF离群点检测的异常检测算法与面向数据空间的最大似然值修复算法,并利用Spark集群对算法进行了并行化的优化处理。最后,开发实现了数据空间数据采集存储与可用性管控系统。基于Flink、Hudi实现了实时与离线采集模块,基于HDFS实现了分布式数据存储模块,并在构建的存储模块上集成了Spark、Flink等多个计算引擎。同时,采用前后端分离的开发模式,后端基于Spring Boot技术架构,前端基于Vue+Element UI开发框架,使用IDEA、VSCode等集成开发工具,完成B/S模式下的数据空间数据采集存储与可用性管控系统开发,并基于数据空间业务数据对采集存储基础架构与数据异常检测与修复算法进行了测试与分析。论文提出了数据空间数据采集存储与可用性管控解决方案,缓解了数据空间在大数据分析上面临的各种问题,但从长远的发展来看,本文依旧存在一些可扩展的地方:简化Hudi表的管理方式、扩展流转数据可用性量化指标以及提升异常检测与修复算法准确性,未来可以在以上方面逐步完善。
【Abstract】 Modern society’s high-speed development of information technology has made networks and software ubiquitous,and communication between people has brought about a large amount of data transmission and circulation.Data has become the indispensable “blood” of today’s society.With the surge in data volume,data processing technology on the automotive industry chain has encountered huge challenges,and the ability to ensure data quality has serious problems.This has triggered a series of severe tests on large-scale data processing platforms in the automotive industry chain in terms of data collection,storage,and availability.Therefore,this thesis focuses on the data space of the automotive industry chain and studies the technology of data collection,storage,and availability control in the data space,aiming to strengthen the big data collection,storage,and management capabilities of the data space and improve the quality of data space intelligent services.This thesis first analyzes the data collection,storage and availability issues and corresponding control requirements of the data space,summarizes the shortcomings of existing solutions,and proposes the overall architecture and various functional processes of data collection,storage and availability control of the data space.It also designs conceptual models,logical models and data tables.Based on Hadoop’s big data processing technology,this thesis proposes a distributed data collection and storage infrastructure that is compatible with the data space and further constructs a data availability control model for the data space.For the infrastructure,it enhances the real-time and offline collection capabilities of the data space based on Flink+Hudi,and implements the storage capability of massive data and multi-computing engine support based on HDFS+Hive+Hudi.For the availability control model,it proposes a quantification index for the availability of dynamic circulating data from the perspective of dynamic circulating data and designs a calculation method for quantification index.From the perspective of static stored data,it proposes an availability improvement strategy.Among them,for the availability improvement strategy,it designs an abnormal detection algorithm based on constraint conditions and LOF outlier detection and a maximum likelihood value repair algorithm oriented to the data space.The algorithm is optimized by parallelization using Spark clusters.Finally,a data collection,storage and availability control system for the data space is developed.Based on Flink and Hudi,real-time and offline collection modules are implemented.Based on HDFS,a distributed data storage module is implemented.Multiple computing engines such as Spark and Flink are integrated on the constructed storage module.At the same time,a front-end separation development mode is adopted.The backend is based on Spring Boot technology architecture,and the front-end is based on Vue+Element UI development framework.Development tools such as IDEA and VSCode are used to complete B/S mode under development.The development of data space data collection storage and availability control system.Based on business data in the data space,tests and analyses were carried out on the designed data space collection storage infrastructure and data anomaly detection and repair algorithms.The thesis proposes a solution for data collection,storage and availability control in the data space to alleviate various problems faced by data circulation and analysis in the data space.However,from a long-term perspective,there are still some areas that can be expanded upon such as resource isolation for real-time collection tasks,evaluation of quantified indicators for circulating data flow rate,expansion of types of abnormal detection algorithms for circulating data flow rate evaluation,expansion of types of abnormal detection algorithms for circulating data flow rate evaluation as well as improving accuracy of repair algorithms for future improvements.
- 【网络出版投稿人】 西南交通大学 【网络出版年期】2024年 12期
- 【分类号】TP311.13;TP333