节点文献

基于数据仓库的青海省统计信息决策分析支持系统设计与实现

【作者】 张伟

【导师】 华庆一;

【作者基本信息】 西北大学 , 计算机软件与理论, 2011, 硕士

【摘要】 面对现在海量的信息数据,如何有效的收集、选择和使用有效数据,更重要的是如何帮助用户在众多的信息中发现它们之间的关系和新的概念,使之能够自动化,并且作用于决策支持是当今信息技术中的一个热点问题。而数据挖掘技术和数据仓库技术是近几年发展起来的有关数据库和人工智能的新技术,它可以通过对大量数据进行建模、分类与聚类等操作,多角度多方位地分析统计数据,发现数据之间的联系。而且在统计工作中,对数据进行统计,并利用数据及其关系来进行决策分析支持,是其工作的一项重要内容,而且统计数据的海量化、零散化是其数据的一个特点,因此将数据挖掘和数据仓库技术与统计行业相结合是一个势在必行的趋势。本文结合“青海省统计信息决策分析支持系统”的开发,对此进行了讨论。根据用户提出的灵活统计、发现数据中的规律、直观展现数据的要求,应用OLAP、决策树等概念描述数据仓库技术,尝试从数据挖掘和数据仓库的角度对统计学进行应用性研究。本文首先介绍系统选题的背景、意义及设计原则,着重归纳国内外数据挖掘和数据仓库领域的专家、学者在数据挖掘这一领域的探索和取得的成果,作为论文的基础。然后从统计行业对数据的要求和具体的数据形式做了基本情况分析、分类分析、动态分析、关联分析、特征值分析等陈述,以明确本文的撰写目的。其次,本文从青海省统计局的统计需求和结构方面出发,应用了统计学的方法明确了挖掘思路及使用的技术,应用数据仓库中的相关技术进行了系统设计,使数据仓库技术在统计行业的应用得以尝试性的实现,对大量原始数据进行了挖掘实现,通过决策树分类方法建立了模型,进一步实现挖掘任务。再者本文对使用数据挖掘技术和数据仓库技术实现的青海省统计信息的数据挖掘与分析系统进行了框架式的介绍,以明证实证部分的设计实现的可行性。最后对本文做了全面的总结,并提出了进一步的努力方向。

【Abstract】 Nowadays, in the face of massive information data, how to effectively collect, choose and make use of effective data, the importantance is how to help the users to figure out the relationships and put forward new concepts in numerous information, making them to be automated, and working on decision support is a hot issue in today’s information technology. However, data mining technology and data warehouse technology have been developed in recent years, relevant to database and artificial intelligence, which can find out useful data from a number of basic data and the relationships between them. Data magnanimity and fragmentation is one of the features in statistical industry and the combination of data mining and data warehouse technology with this industry is just an imperative trend. This paper discussed it with Qinghai Province Statistical Information Decision Analysis Support System Design And Realization.Accord to user’s request of Flexible query, finding rule through data, data visualizing, using OLAP, decision tree, such as concept description data warehouse technology. This paper attempts to provide an applied research on statistics from the perspective of data mining and data warehouse.Firstly, this thesis introduces the background of the topic, significance and design principles, focusing on summarizing the results obtained by the experts and scholars in the field of data mining and data warehouse at home and abroad as the theoretical basis. In order to identify the purpose of this thesis, it also describes the basic situation analysis, classification analysis, dynamic analysis, correlation analysis, eigenvalue analysis and the other statements from the specific data requirements and data forms in statistical industry.Secondly, in the demonstrative part of this thesis, I have applied practical design to the statistical demands and structure of the Qinghai Provincial Bureau of Statistics with the related technology in data warehouse, trying to achieve the application of data warehouse technology in the statistical industry. In the process of description, I used the statistical way of thinking to decide the mining ideas and techniques to be used, and a large number of raw data to implement data mining. Through the methods of decision tree classification, a model is established for further mining tasks.And then this thesis reveals a frame presentation on Qinghai statistics data mining and analysis system achieved by the use of data mining and data warehouse technology, to prove the feasibility of the design and implementation in the last part.Finally, there is a comprehensive summary in the thesis and the direction of further efforts have been proposed as well.

  • 【网络出版投稿人】 西北大学
  • 【网络出版年期】2012年 05期
  • 【分类号】TP311.13
  • 【被引频次】1
  • 【下载频次】131
节点文献中: