节点文献

基于模糊系统和遗传算法的数据挖掘技术的研究

Research of Data Mining Techniques Based on Fuzzy System and Genetic Algorithm

【作者】 张巍

【导师】 侯国莲;

【作者基本信息】 华北电力大学(北京) , 控制理论与控制工程, 2003, 硕士

【摘要】 数据挖掘是分析大规模数据集的有效方法。由于数据内在的不精确性和多属性之间的复杂性,有时已有的方法就失效了,而软计算技术在这两方面有着独到的优势,所以以软计算技术为手段研究新的数据挖掘方法具有重要的意义。本文力图将模糊系统和遗传算法结合起来,发挥各自的优势,研究解决数据挖掘的两类基本问题--预测和分类的新方法。本文的工作主要体现在以下几方面:1)回顾了数据挖掘的历史背景,总结了数据挖掘的定义,数据挖掘的过程,挖掘的六种知识及其主要方法。2)概要的介绍了本文研究内容的理论基础:模糊系统和遗传算法。3)针对预测问题,本文提出了一种由模糊规则构成的预测模型,并用遗传算法来确定模糊系统的参数,最后以某气象台的历史资料为依据预测降水量,作为对比,本文同时给出回归分析方法得出的结果,对比表明本文的方法得出的结果要优于回归分析方法得出的结果。4)针对分类问题,本文提出了一种由模糊规则构成的分类器,每条模糊规则由研究对象一个属性、归属类和影响因子组成。遗传算法被用来寻找最优模糊规则集。在算例中,本文将此模型用于机器学习领域的一个经典数据集,得到了规模小、分类准确度高的分类器

【Abstract】 Data mining is a set of methods efficient in analyzing large data sets, however for the inherent uncertainty and complex of data and the attributes, some methods show their inability in some cases. Soft computing is good at dealing with such dilemma, therefore it is valuable to study data mining techniques in the frame of soft computing. In this paper, the combination of fuzzy model and genetic algorithm is presented to mine two kinds of knowledge: prediction and classification. The following points are concentrated in this paper:1) The background of data mining’s emergence, the definition and the process of data mining, the mined knowledge and corresponding methods are presented. 2) As the basis of following study, the theory on fuzzy model and genetic algorithm is briefly introduced.3) For prediction, a method is proposed which is based on the model of fuzzy system and optimized by improved genetic algorithm. As an instance, the historical data of some forecast station is analyzed using the method to predict the amount of precipitation, the contrast with the result from multivariable regression shows that the method has better performance.4) For classification, a classifier composed of fuzzy rules is given, the fuzzy rules are simple consisting of an attribute, the corresponding class and an influence factor. Genetic algorithm is used to find the best subset of fuzzy rules. In the instance, the classical data set of Iris is studied, which is studied widely in machine learning domain, as the result, a subset with less rules and more accuracy is obtained

  • 【分类号】TP311.13
  • 【被引频次】11
  • 【下载频次】401
节点文献中: 

本文链接的文献网络图示:

本文的引文网络