节点文献

TGFCM:基于模糊聚类的中文文本挖掘的新方法

TGFCM: A Novel Approach of Chinese Text Mining Based on Fuzzy Clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 耿新青王正欧

【Author】 GENG Xinqing,WANG Zheng’ou(Institute of Systems Engineering,Tianjin University,Tianjin 300072)

【机构】 天津大学系统工程研究所天津大学系统工程研究所 天津300072天津300072

【摘要】 提出一种新的动态模糊聚类的方法,针对传统的模糊聚类需要预先确定聚类数的问题,提出采用动态自组织映射神经网络来确定聚类数,并通过文本向量空间模型和TF?IDF方法来确定文本的特征向量,再将动态自组织映射神经网络得到的聚类数,用模糊C均值算法(FCM)函数处理,得到聚类的结果。该算法同仅用动态自组织映射神经网络算法的运行结果相比,具有运行聚类结果精度高的优点,模糊聚类更适合处理语义的多样性和文本归属的模糊性,实验验证了算法的有效性。

【Abstract】 A novel approach is presented.The main defect of traditional methods of fuzzy clustering is to known the number of clustering in advance.This paper applies the dynamic self-organizing maps algorithm to determining the number of clustering.The text eigenvector is acquired based on the vector space model(VSM) and TF?IDF method.The result of clustering is attained by fuzzy C mean algorithm(FCM).The number of clustering acquired by the dynamic self-organizing maps is introduced into the fuzzy C mean algorithm(FCM).Compared to the dynamic self-organizing maps algorithm,the present algorithm possesses higher precision.The fuzzy clustering is suitable for dealing with the semantic variety and complexity.The example demonstrates the effectiveness of the present algorithm.

【基金】 国家自然科学基金资助项目(60275020)
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2006年05期
  • 【分类号】TP391.1
  • 【被引频次】4
  • 【下载频次】271
节点文献中: 

本文链接的文献网络图示:

本文的引文网络