节点文献

数据导入和预处理系统设计与实现

Design and Implementation of Data Import and Preprocessing System

【作者】 杨毅

【导师】 李明楚;

【作者基本信息】 大连理工大学 , 软件工程(专业学位), 2017, 硕士

【摘要】 传统数据仓库随着Hadoop技术的发展受到巨大挑战,Hadoop从最初解决海量数据的存储难题,到现在被越来越多的企业用来解决大数据处理问题,其应用广泛性越来越高。本文主要研究基于Hadoop系统对传统数据库数据和文本数据进行迁移,帮助传统数据仓库解决在大数据存储处理等方面遇到的难题,同时依靠Hadoop的扩展性提升数据存储和处理的性能。论文中系统根据现今传统数据仓库的应用情况及Hadoop大数据平台的前景预测,针对传统数据仓库已无法满足用户需求的问题,设计出传统数据仓库与基于Hadoop的hdfs文件系统协作进行数据存储与处理的架构,同时解决企业用户数据控制权限的要求。系统分为四个部分,数据管理、数据预处理、系统管理和发布管理提供从数据导入到数据控制,数据预处理最终实现数据发布共享的功能。系统的主要功能是采集数据和对采集到的数据进行预处理,系统设计成能够对多种类型的数据进行采集和预处理,同时系统能够实现很好的扩展功能,为系统中增加机器学习算法节点对数据进一步挖掘处理提供了可能。系统采用当下流行的Hadoop基本架构,同时结合Haddoop生态圈中的数据仓库Hive和数据迁移工具Sqoop进行数据的迁移和处理。在一定程度上能够满足企业的基本需求。系统以Web系统的方式实现,方便用户使用,在实现Web系统时采用成熟的ssm框架进行开发,保证系统的稳定性。系统从企业的实际需求出发,同时充分考虑传统数据库在企业中的应用,设计实现基于Hadoop的数据管理平台原型,为企业提供实际应用指导。本论文从系统实现的背景、系统系统需求、系统设计、系统实现以及系统测试五大模块对系统进行了全面详细的论述,全面阐述了系统实现的意义,有一定的实际应用指导意义。

【Abstract】 With the development of hadoop technology,from the initial Google,Facebook and other companies to solve the massive data storage problems,and now more and more enterprises to deal with large data,enterprises have built a good traditional data warehouse status has been challenged.This paper focuses on how hadoop works with traditional data warehouses,how to carry out transmission,storage,and processing.Based on the traditional data warehouse has been provided on the basis of hadoop support to make up for the traditional data warehouse in the massive data processing,storage and other deficiencies,can also rely on Hadoop’s horizontal scalability to break through a single node of the traditional data warehouse in storage and computing power The bottleneck.Based on the application of traditional data warehouse and the prospect of hadoop large data platform,this paper designs the traditional data warehouse and the architecture of data storage and processing based on hadoop hdfs file system for the problem that traditional data warehouse can not meet the needs of users.,While addressing the enterprise user data control permissions requirements.The system is divided into four parts,data management,data preprocessing,system management and release management from data import to data control,data preprocessing and ultimately data publishing and sharing functions.The main function of the system is to collect data and to pre-process the collected data.The system is designed to collect and preprocess various types of data.At the same time,the system can achieve very good extended functions,adding machine learning algorithms Nodes to further dig data processing possible.The system uses the current popular Hadoop infrastructure and migrates and processes data with Hive,a data warehouse in the Haddoop ecosystem,and Sqoop,a data migration tool.To a certain extent,to meet the basic needs of enterprises.System to achieve the Web system,user-friendly,in the realization of the Web system using mature ssm framework for development,to ensure system stability.This system is based on the demand of large-scale enterprise platform,but also takes into account the reuse of traditional data warehouse,the cooperation between the two,and finally realize the system prototype to provide guidance for the practical application of the enterprise.This dissertation discusses the system in detail from the background of the system realization,the system system requirements,the system design,the system realization and the system test.It expounds the significance of the system implementation comprehensively,and has certain practical significan.

【关键词】 Hadoop数据仓库数据预处理
【Key words】 Hadoopdata warehousedata preprocessing
  • 【分类号】TP311.13
  • 【被引频次】2
  • 【下载频次】136
节点文献中: