节点文献
基于粗糙集的相对属性约简算法及决策方法研究
Research on Relative-Attribute Reduction Algorithm and Decision-Making Method Based on Rough Set
【作者】 张国军;
【导师】 卢炎生;
【作者基本信息】 华中科技大学 , 计算机软件与理论, 2010, 博士
【摘要】 粗糙集理论和方法是一种能有效的分析和处理不一致、不精确、不完备等各种信息的数据分析工具。该理论和方法已经在模式识别、机器学习、决策支持、知识发现、预测建模等领域得到成功的应用。相对属性约简算法和决策方法是粗糙集理论和应用的关键技术之一,也是知识发现和决策的重要研究课题,并已成为一个备受关注的研究热点。围绕粗糙集相对属性约简和决策方法中的相对属性约简、决策规则获取、基于粗糙集的决策方法以及其原型系统等四个重要问题,从六个方面开展研究工作。它们分别是基于贪心策略的相对属性约简、基于核的自顶向下剪枝的相对属性约简、基于贪心策略的分类规则获取、基于粗糙集的支持向量决策方法、基于粗糙集的最优支持向量决策方法和基于粗糙集的支持向量集成决策模型。针对寻求单个相对属性约简的问题,基于贪心策略的相对属性约简算法以条件属性的分类能力作为启发信息,是解决相对属性约简的一种有效算法,该算法的实现比较直观。当决策信息系统中包含有大量对象时,该算法有效节约了存储空间,适合大规模数据集上的计算,而且在该相对属性约简方法中,仅仅考虑了属性分类能力大小,倾向于选择分类能力强的条件属性加入到相对属性约简中,根据分类目标,这种倾向是合理的。当一个决策信息系统包含相当多的属性和大量的纪录时,如何从决策信息系统中获取包含最少条件属性的相对属性约简和获取所有相对属性约简的集合是一个值得研究的课题。基于核的自顶向下剪枝的相对属性约简算法是解决该问题的一种可行算法,实验结果表明基于核的自顶向下剪枝的相对属性约简算法的可行性和有效性。在决策规则的获取方面,根据决策规则的不同度量,从不同的角度获取决策规则,可获得基于贪心策略的分类一致性规则获取算法和不一致性分类规则获取算法。这些算法根据属性的决策能力的大小作为启发式知识来指导这一属性值约简过程的进行,不但获取的规则通常较短,而且有较强的分类预测性能,既提高了运行速度,又节约了存储空间。随着决策信息系统的数据量的增加,粗糙理论分类的容错能力与泛化能力较弱等缺点也突现出来,因此,如何提高决策信息系统的容错能力与泛化能力是一个值得研究的课题。从不同的角度,可获得基于粗糙集和支持向量优点的三种决策方法。通过对比实验,结果表明相关算法有较高的容错能力与泛化能力。基于前面的研究结果以及有关技术,设计并实现了一个基于粗糙集的相对属性约简算法和决策方法的原型系统。与同类系统相比,该系统在设计实现上具有一定的独到之处,具有较高质量的知识发现和决策结果,并具有较好的容错能力与泛化能力,还具有较强的鲁棒性。
【Abstract】 Rough set is an excellent data analysis tool to process inconsistent, inaccurate, incomplete information and so on. Its theories and methods have been applied to many areas successfully including pattern recognition, machine learning, decision support, knowledge discovery, and forecast modeling and so on. The relative-attribute reduction algorithm and decision-making method are key technologies of rough set theory and application and becoming a research focus of concern.Surrounding the four key problems of relative-attribute reduction and decision-making method, i.e. relative-attribute reduction, rules of acquisition, decision method based on rough set and its prototype system, the following six works have been done: greedy relative-attribute reduction algorithm, top-down pruning relative-attribute reduction algorithm based on core, greedy decision rules of acquisition, support vector decision-making methods based on rough set, optimal support vector decision-making methods based on rough set, and decision-making model of support vector ensembles based on rough set.For the problem of questing single relative-attribute reduct, the greedy relative-attribute reduction algorithm, which takes classification ability of condition attributes for heuristic information, can be an effective algorithm to achieve more intuitive to solve the above. When the decision-making information system contains a large number of objects, the algorithm can save a lot of storage space and is suitable for large-scale data sets on the calculation. It tends to choose the condition attributes of higher classification ability added to the relative attribute reduct. According reduction goal, this tendency is reasonable.When a decision-making information system consists of a considerable number of attributes and a large number of records, how to get the relative-attribute reduct including least condition attributes or the whole set of all relative-attribute reducts from the decision-making information systems, is a worthy subject of research. Top-down pruning relative-attribute reduction algorithm based on core is a viable solution to the problem. Experimental results show that top-down pruning relative-attribute reduction algorithm based on core is feasible and effective.According to different measure of decision-making rules from different aspects, the greedy methods obtaining decision-making rules of consistency or inconsistency are proposed. The methods, which take classification ability of condition attributes for heuristic knowledge to guide the attribute value reduction process, not only get the shorter rules of the strong performance in the classification of forecasts, but also improve the speed and save the storage space.By following the increasing of the amount of data, the weakness of fault tolerance and generalization capability is to the fore in rough set theory. Therefore, it is a worthy subject of study how to improve fault-tolerant ability and generalization ability of the decision-making. Three kinds of support vector decision-making methods based on rough sets are proposed from different aspects. By comparing the other methods, experimental results show that our proposed algorithms have higher fault-tolerant ability and generalization ability.Based on the above research findings, as well as the related technology, designed and implemented a prototype system based on relative-attribute reduction algorithms and decision-making methods. Compared with similar systems, this prototype system is designed to achieve that has certain uniqueness, with higher quality results of knowledge discovery and decision-making, and has good fault-tolerant ability and generalization ability, but also has strong robustness.
【Key words】 Rough Set; Greedy Strategy; Top-Down Pruning; Relative-Attribute Reduction; Rules of Acquisition; Support Vector; Decision-Making Ensemble Model;