节点文献

基于神经网络等技术的数据与文本聚分类研究

Clustering and Classification of Data and Text Using such Technologies as Neural Network

【作者】 钱晓东

【导师】 王正欧;

【作者基本信息】 天津大学 , 管理科学与工程, 2005, 博士

【摘要】 聚类和分类技术是数据挖掘中最有价值的技术之一,而软计算中的神经网络是聚分类中的主要技术之一。自适应谐振神经网络(Adaptive Resonance Theory:ART)不仅参考人脑神经元互连的物理模型,而且也借鉴人脑的学习机理,具备数据聚类的良好特性,目前国内外研究尚较处于发展阶段。文本挖掘中文本向量集往往表示为正交的高维空间,因而带来计算瓶颈和与实际应用背景不吻合的情况,研究特性良好的降维算法、现有空间的改进等都存在很大的发展余地。本论文提出了四种基于ART2神经网络的用于数据聚类的改进算法,克服了经典ART2神经网络输出无层次结构的缺点,均可形成动态的层次聚类结果,同时降低了警戒参数主观设置的要求。基于模、相位、空间密度的改进ART2算法1还克服了经典ART2算法警戒参数全局化、聚类与模无关的缺点,其通过按模和相位的综合评价,依据先前循环形成类别中的输入向量个数分类别修正警戒参数以实现按空间密度局部化警戒参数,在借鉴以前神经网络训练结果的基础上进行聚类;基于凝聚和迭代思想的改进ART2算法2通过迭代在人工交互下达到合理聚类结果,并计算出合理聚类结果所需的警戒参数范围值;迭代以及迭代中神经网络的输出都体现出有序的自组织特征,网络训练时间代价也在迭代中迅速下降;基于Hebb规则和泄漏竞争的改进ART2算法3借鉴了Hebb规则和泄漏竞争的思想,允许多个神经元获胜并计算获胜神经元之间的相关性;基于Hebb规则和冗余神经元思想的改进ART2算法4克服了过分依赖获胜神经元信息等不足,通过在竞争过程中同时考虑获胜神经元和其它神经元的信息以及Hebb规则来实现通过单个ART神经网络的层次聚类结果。本论文提出了一种基于随机映射的文本降维算法,在可控、低代价地充分逼近原始空间相似度计算结果和分类结果的情况下降低文本向量空间维数。在此基础上本论文还提出了一种基于随机映射的加速隐含语义索引算法,此加速算法将随机映射和隐含语义索引相结合,既可有效可控地降低空间维数,又可凸现语义联系,使得其用于分类算法在文本高维环境中具备实时性和高分类准确率。此外本论文提出了一种基于模式聚合和各维不同权重的改进KNN文本分类算法,在数据分析的基础上提出优化的模式聚合方法,并利用神经网络计算空间各维不同权重以克服VSM空间各维权重相等的缺点,可以在降低时间和空间复杂度的基础上,提高KNN算法的文本分类准确度。

【Abstract】 Clustering and classification are one of the most valuable technologies in datamining, and the neural network in soft calculation is one of the main technologies ofclustering and classification. Adaptive Resonance Theory(ART) neural network notonly refers to the physical connection model of human brain neuron, and also to thestudying mechanism of human brain, therefore has the good feature of data clustering.The researches of ART are now still in the beginning phase. In text mining, text vectorset is usually expressed as the high dimensional orthogonal space, therefore, it bringsthe calculation bottleneck and inconsistence with the factual application background.So, the researches of good dimension decreasing algorithm and the improvement ofcurrent space have a lot of developing space.This dissertation presents 4 kinds of improved algorithms based on ART2 neuralnetwork for data clustering. All these improved algorithms overcome the classicalART2’s shortcoming such as output without hierarchical network and form thedynamic hierarchical clustering results, and meanwhile decrease the requirements ofvigilance parameter’ subjective configuration.The improved algorithm 1 of ART2, which is based on the integration of modul,phase and space density, also overcomes the classical ART2’s shortcomings includingvigilance parameter globalization and clustering independence with mode, andclusters by the comprehensive comments of modul and phrase and together with thereference to the previous training result of the neural network. It adjusts the vigilanceparameter according to the number of the input vectors of the classes generated inprevious cycle to realize the vigilance parameter localization based on the spacedensity.The improved arithmetic 2 of ART2 based on the agglomeration and iterationachieves the reasonable clustering result by the manual interaction through iterativemethods, and calculates the required vigilance parameters range necessary to thereasonable clustering result;all the iterative process and the output of neural networkin the iteration process exhibit orderly self-organization feature, and the networktraining time also rapidly decreases in the iterative process.The improved algorithm 3 of ART2 based on Hebb rule and the leakedcompetition allows the multi-neurons to be the winners and calculates the correlationamong the winning neurons.The improved algorithm 4 of ART2 based on Hebb rule and redundant neuronovercomes the shortcomings including relying on winning neuron too much and so on,and it implements hierarchical clustering result using single ART neural network bythe method of considering both the winning neuron and other neuros’ information andtogether with Hebb rule.This dissertation also presents a text dimension-reduction algorithm based onrandom mapping(RM), under the conditions of controllable and low cost, andapproximating sufficiently to the calculation and classification results of the originalspace, it can greatly decrease the dimension of the text vector space. On the basis ofthis algorithm this dissertation presents an accelerated latent semantic index(LSI)algorithm that is based on the combination of RM algorithm and latent semantic index,the accelerated LSI algorithm can efficiently reduce the dimension of the space andalso emphasize the semantic relationship, therefore it makes the classificationalgorithms have real-time and better classification accuracy in high-dimension textenvironment.In addition, this dissertation carries out a improved KNN text classificationalgorithm based on pattern aggregation and different weights of each dimension. Onthe basis of data analysis the improved pattern aggregation method is presented. Andthe neural network is used to calculate the weight of each dimension of VSM modelsto overcome such a shortcoming of VSM space as possessing the same weight byeach dimension. Therefore it can increase the text classification precision of theimproved KNN algorithm on the basis of decreasing of the complexity of time andspace.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2006年 07期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络