节点文献
聚类和神经网络算法研究及其在电信业客户消费模式中的应用
Application in Customers Consumption Model of Telecom Based on Clustering Algorithm and Neural Network
【作者】 洪晶;
【导师】 柳炳祥;
【作者基本信息】 景德镇陶瓷学院 , 机械设计及理论, 2007, 硕士
【摘要】 数据挖掘(Data Mining),又称知识发现(KDD),是从大量的、不完全的、有噪声的、模糊的、随机的实际应用数据中,提取隐含在其中的、人们事先不知道的、但又是潜在有用的信息和知识的过程。它是一门新兴的交叉学科,汇集了来自机器学习、模式识别、数据库、统计学、人工智能等各领域的研究成果。其中聚类和神经网络是数据挖掘中最常用的两种算法。论文主要研究了K-means聚类算法和BP神经网络,并将它们结合起来应用于电信业客户消费模式的研究。聚类是一个将数据集划分为若干组(class)或类(cluster)的过程,并使得同一个组内的数据对象具有较高的相似度,而不同组中的数据对象则是不相似的。K-means算法是聚类算法中主要算法之一,是一种基于划分的聚类算法。该算法随机选取K(K为聚类数)个点作为初始聚类中心,通过一个迭代过程完成聚类。如果初始聚类中心选取不合理,就会误导聚类过程,得到一个不合理的聚类结果。论文对K-means算法中初值的选取方法进行了分析和研究,提出了一种新的选取初始聚类中心的方法,提高了聚类准确率。此外,BP算法作为最常用的神经网络算法也是论文研究的重点之一。虽然BP网络预测模型结果不错,但是单纯的BP算法自身存在着一些不足:(1)易陷入局部极值;(2)遗忘已学样本的趋势;(3)学习效率不高,收敛速度慢等。论文将模拟退火(SA)算法来优化BP网络,很好地避免了BP算法的收敛速度慢,易陷入局部极值点的问题。通过实验分析,取得了很好的预测效果。因此,论文首先利用统计学相关分析方法去除建模中的冗余字段,然后建立了一种基于聚类分析和神经网络算法的分类预测模型,并将所建立的分类模型应用到电信业客户消费模式中去,预测出每一位客户最终所属的消费模式类别,能够帮助客户服务人员按照每一类客户群体消费行为的特点提供相应的服务和采取针对性的营销策略,从而根据潜在客户消费模式,对现有客户提供更好的服务,同时发掘出潜在客户及需求,最终为公司带来更大的利润。
【Abstract】 Data mining, also referred to Knowledge Discovery in Database (KDD), means a process of nontrivial extraction of implicit, previously unknown and potentially useful information (such as knowledge rules, constraints, regularities) from data in database. It is an emerging interdisciplinary studies, collects kinds of research results such as Machine Learning, Pattern Recognition, Database, Statistics, Artificial Intelligence and so on. Clustering and Neural Network are two of the most common algorithms in Data Mining. This paper has mainly studied the K-means clustering algorithm and the BP neural network, and unifies them applies in the research of telecommunication customer consumption pattern.The clustering is the process which divides the data set into certain group of (class) or the kind of (cluster), and enables the data object in the identical group to be similar, but the different groups data object are in a big difference. The K-means algorithm is one of.main clustering algorithms. It is a kind of the clustering algorithms based on partitioning methods. This algorithm selects K (K is stochastically cluster number) the spots as the center of the initial cluster, completes the cluster through an iterative process. If the initial cluster center is selected unreasonably, which will mislead the cluster process, and come into an unreasonable result. This paper analyses and studies the selection methods of initial numbers in K-means Clustering, proposed one new method to select initial cluster center, and enhanced the cluster rate of accuracy.In addition, the BP algorithm, which is the most common neural network algorithm, also is one of key research in this paper. Although the result of the BP. network forecast model is good, but pure BP algorithm has some insufficiencies: (1) easy to fall into the partial extreme value; (2) the tendency in forgetting studied samples; (3) study efficiency is not high, convergence rate slow and so on. This paper will simulate the annealing (SA) algorithm to optimize the BP network, avoided the BP algorithm convergence rate being slow well, easy to fall into the partial extreme point question. It analyzed through the experiment, and has obtained better-forecast effect.Therefore, this paper first uses in statistics correlation analysis method to eliminate the redundant field of the model, and then establishes a classification forecast model based on Clustering analysis and Neural Network algorithms, and will establish the classified model to apply in the telecommunication customer consumption pattern, confirms the algorithm the feasibility and the validity. Thus can help the customer servicers to provide the corresponding service according to the customer consumer behavior characteristic and to adopt the pointed marketing strategy. According to the latent customer expense pattern, it may provide a better service to the existing customers, and select the latent customer and their demand, finally brings a bigger profit for the company.
【Key words】 Data mining; Clustering; K-means Algorithm; Neural Network; BP Algorithm; customers consumption pattern;
- 【网络出版投稿人】 景德镇陶瓷学院 【网络出版年期】2008年 02期
- 【分类号】TP183;F626
- 【被引频次】9
- 【下载频次】618