节点文献

基于遗传算法的聚类数据挖掘及其在销售系统中的应用

Genetic Algorithm Based Clustering Data Mining and Application in Sale System

【作者】 王武龙

【导师】 黄明;

【作者基本信息】 大连铁道学院 , 交通信息工程及控制, 2002, 硕士

【摘要】 随着信息化时代的到来,信息资源的经济价值和社会价值越来越明显。从大量的数据资料中发现有价值的信息或知识,达到为决策服务的目的,成为非常艰巨的任务。我们需要从大型数据库或数据仓库中提取人们感兴趣的知识,这些知识是隐含的、事先未知的潜在有用信息。数据挖掘方法的提出,让人们有能力最终认识数据的真正价值,即蕴藏在数据中的信息和知识。数据挖掘是目前国际上数据库和信息决策领域的最前沿研究方向之一。 数据挖掘具有广泛的研究领域,其中聚类算法挖掘的研究开展得比较积极,本文主要围绕遗传算法对数据挖掘理论和方法进行了以下几方面的工作: 首先归纳了数据挖掘技术的总体研究情况,包括数据挖掘的定义、与其他学科的关系、挖掘的主要过程、分类和主要技术手段。重点探讨了数据仓库和数据挖掘的关系。数据仓库作为一种新型的数据的存取地,为数据挖掘提供了新的支持平台。因其内在的对决策的支持能力,为数据挖掘开辟了新的空间。 其次深入研究了数据挖掘领域中的一个重要研究方向——聚类。聚类技术在统计数据分析、模式识别、图像处理等领域有广泛应用迄今为止人们提出了许多用于大规模数据库的聚类算法。本文对遗传算法进行深入研究,对遗传算法进行了优化改进,提出一种高效的基于遗传算法的聚类挖掘。聚类的数据挖掘主要挑战性在于数据量巨大且必须全局遍历,因此应用遗传算法的效率是很关键。 然后构建销售系统数据仓库初型。完成大钢集团销售系统的设计与开发,包括初步设计、详细设计以及软件开发。大连钢铁集团CIMS工程销售系统主要包括订货子系统、发货子系统、价格子系统和资金子系统。本文的基于遗传算法的数据挖掘主要以大连钢铁集团CIMS工程订 大连铁道学院工学硕士学位论文 货于系统为背景。开发的系统现己正常运行。 最后在研究遗传算法理论与应用的基础上,将改进后的算法应用于 大钢销售系统。

【Abstract】 With the development of the information technology,the economical and social value of information resource is more and more important. In order for decision,it is the hard work that the valuable information or knowledge is discovered from numerous data. We need discover the interesting knowledge from the data warehouse. This knowledge is useful,but it cannot been predict. Because the method of Data Mining is put on,the people can realize the real value of data,which mean the information or the knowledge hide in data. Data Mining is one of the international advanced directions in the field of database and information decision.Data Mining has many research fields. The research of clustering algorithm is advanced. This dissertation is mainly on the theory and method of Data Mining and genetic algorithms based clustering algorithm. All the work can be concluded as follows:Firstly,induces the research work and development of Data Mining,including the Data Mining definition,the relationships with other academic fields,the imperative processes and its classifications. The principal techniques used in the Data Mining are surveyed also. The research of Data Mining and data warehouse relationship is the main work too. Data warehouse is the new saving data place for Data Mining,because it supportsdecision.Secondly,clustering is a promising application area for many fields including data mining,statistical data analysis,pattern recognition,image processing,etc. Many clustering algorithms have been developed. In this dissertation,on the base of clustering algorithm and genetic algorithms research,this dissertation optimized and improved the genetic algorithms,and presents an efficient algorithm. In the field of Data Mining research,genetic algorithm is a promising algorithm. The main challenge and key problem that clustering algorithm is applied to data mining is enormous data and to scan it,so efficiency is very important.Thirdly,build up the system prototype of data warehouse.Finish the system analysis,design and development,including the initiative design,detailed design and software design. This dissertation is under the environment of DG-CIMS including Order subsystem,Consign subsystem,price subsystem,and funds subsystem mainly. The system that has been developed is running successfully for 10 months.Lastly,the new algorithm gives some experimental results in DG-CIMS. On the base of the Data Mining and genetic algorithm based clustering algorithm research,the improved algorithm verifies the performance in DG-CIMS.

  • 【分类号】TP311.12
  • 【下载频次】279
节点文献中: