节点文献
基于密度的数据流聚类挖掘算法
Based Density Data Stream Cluster Mining Algorithm
【作者】 王延明;
【导师】 寒枫;
【作者基本信息】 华北电力大学(河北) , 计算机应用技术, 2007, 硕士
【摘要】 近年来,越来越多的应用产生数据流,它是连续的、有序的、快速变化的、海量的数据。流数据不同于传统的存储在磁盘上的静态的数据,而是一类新的数据对象。当前在数据挖掘领域当中,数据流已经成为一个研究热点。数据流聚类分析成为聚类研究的一个重要方向。本文首先介绍了数据挖掘的相关知识,并对数据流挖掘进行了论述,然后建立了一个数据流聚类算法DSCluster,相对于别的双层数据流聚类算法,除了在空间和时间效率上保持一致外,本算法最大的特点就是可以对混合属性的数据流进行聚类分析,并且可以更好的检测数据流当中的异常点。接下来通过相关的实验,将DSCluster算法和其他算法对比,显示了DSCluster算法的高效性和先进性。最后对本文的内容进行了总结,并对以后的工作进行了展望。
【Abstract】 Recently, there are more and more applications that are facing the envirnoment of stream data. Stream data is a kind of continuous; ordered, changing fast and huge amount data. It is quite a new object that is different from traditional static data stored on the disk. Currently, data mining in data stream becomes a hot research field. First, we introduce the knowledge of data mining and discuss the data stream mining, then we build a data stream mining algorithm—DSCluster which may cluster and detect outliers in data stream containing both continuous and categorical attributes. Furthermore, the paper reports experiments on real-life datasets and synthetic datasets, the results show that our algorithm can get higher accuracy of clustering within limited memory, and has the good scalability with the quantity and the dimensionality of stream data. Finally, we summarize the content of paper and point out the research emphases for future work.
- 【网络出版投稿人】 华北电力大学(河北) 【网络出版年期】2007年 01期
- 【分类号】TP311.13
- 【下载频次】339