节点文献

数据挖掘方法研究:关联和趋势分析

Research of Data Mining Method: Association and Trend Analysis

【作者】 曾异平

【导师】 朱宏;

【作者基本信息】 电子科技大学 , 运筹学与控制论, 2003, 硕士

【摘要】 本文研究了两类数据挖掘方法。全文分五个部分:引言、数据挖掘方法概述、关联分析方法研究、趋势分析方法研究和结论。在引言中介绍了数据挖掘产生的原因:数据的急剧膨胀和高度时效性与人们得不到科学决策所需要的有效信息和知识之间的矛盾;给出了数据挖掘的发展和演化过程;然后指出了数据挖掘前景,最后叙述了本文所做的全部工作。在第一章数据挖掘方法概述部分,重点阐述了数据挖掘的定义、数据挖掘方法分类、数据挖掘方法研究现状以及数据挖掘和统计学的区别与联系。指出了数据挖掘定义所包括的几层含义:面向真实数据、面向具体问题等;给出了数据挖掘方法的分类,确定了本文研究的两类挖掘方法在整个数据挖掘方法中的地位和作用;从八个方面详细总结了现阶段数据挖掘方法的研究现状;最后对数据挖掘与统计学的关系进行了讨论,指出了数据挖掘与统计学相同之处和本质区别。在第二章关联分析方法研究部分,重点讨论关联分析的经典方法和基于兴趣度的否定关联分析方法。通过一个实例,指出了经典关联分析方法在“支持度-置信度”框架下产生了错误的关联规则;并针对这种情况,提出了基于兴趣度的否定关联分析方法,对所举实例进行分析,表明该方法能挖掘出更加符合实际的、用户感兴趣的否定关联规则。该方法采用卡方统计量作为兴趣度度量,并修改经典关联分析方法:方法,以进行否定关联分析。在第三章趋势分析方法研究部分,通过对交易数据项集进行编码把原始数据转换成整数值随机变量序列,并说明了该序列为马尔可夫链,然后用频率代替转移概率,建立了一个趋势分析的模型。对超市销售数据进行分析表明该方法简单、实用,而且得到一个有趣的结果:顾客对同一产品的不同品牌的选择是没有差别的。在第四章结论部分,对本文在数据挖掘方法上的研究工作进行了总结。

【Abstract】 This dissertation mainly studies two methods of Data Mining and consists of five components: introduction, the summarization of Data Mining, the research of associate analysis, the research of trend analysis and conclusion.In the introduction, the reason of studying Data Mining is given: conflict between the exploring and updating of data with lacking of effective information and knowledge that people need in wise decision. Then the development and evolvement of Data Mining are given. In the end of the introduction the prospect of Data Mining is discussed.In chapter 1, the definition of Data Mining, the category of methods of Data Mining, the front of Data Mining and the deference and correlation of Data Mining and statistics are discussed detailed. First, investigating two necessary definitions in the definition of Data Mining: having the true data and facing the real problem. The result of two classify of methods of Data Mining that are discussed in this paper lead to the important position in the research of Data Mining. Then summarizing the status of research in Data Mining from eight aspects. Last, in discussing the relation of Data Mining and statistics, the sameness and distinguish between them are given.Chapter 2 Mainly discussing the classic method of associate analysis and the method of negative associate analysis that based on interest measurement. Through a case the false rule that concluded from the classic method of associate analysis that based on "support-confidence". While analyses the same case using the method of negative associate analysis that based on interest, we can see that this method can mine more applicable and more interesting data for the user. In order to apply the method of negative associate analysis this method has the as its interest measurement and modify the classic method: Apriori method.Chapter 3 Based on encoding the transaction data item set, the original data is<WP=5>transformed into a series of integral variable and this series is proved to be a Markov chain in theory. And instead of transfer probability, frequency is used in transfer probability matrix. Then analysis for the sale data of supermarket indicates that the method is fine, and a nice result that customs choice is same to different bland of one product is received.In the conclusion part, the author summaries the research of Data Mining ’s method.

  • 【分类号】TP311.13
  • 【被引频次】8
  • 【下载频次】1064
节点文献中: