节点文献

一种自动搜索阈值的中文文本层次聚类方法

A Chinese Text Hierachical Clustering Method Based on Auto-searching Threshold

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 向继荆继武高能

【Author】 Ji Xiang Jiwu Jing Neng Gao (state key laboratory of information security(GUCAS),Beijing 100049)

【机构】 信息安全国家重点实验室(中国科学院研究生院)

【摘要】 文本聚类是分析和处理网络文本的重要手段,文本层次聚类是目前最常用的方法之一。本文通过研究和分析传统的文本层次聚类方法的不足,提出了一种改进的基于阈值自动搜索的方法。该方法利用簇集的相似性分布和最小二乘曲线拟合方法自动发现层次聚类中每次迭代的阈值,同时用固定的两次迭代取代原来的不定次数的多次迭代,避免了由用户来设置聚类参数,提高了聚类的自动性。通过实验结果表明,该方法在聚类准确性上比传统的方法有所提高,而且该方法在孤立点容忍和防止错误扩散方面也有一定的进步。

【Abstract】 Text clustering is an important means for analyzing and processing network texts, and hierachical clustering is one of the most used text clustering methods.After studying and analyzing weakness of traditional text hierachical clustering method, this paper proposed an improved method based on auto-searching threshold.The proposed method utilized cluster set similarity distribution and the least-squares curve fitting technology to search the clustering threshold automatically, and it used a fixed two rounds iteration instead of traditional variable rounds iteration.This method avoided parameter setting by users, so that it improved the automaticity of clustering process.The experiments show that the proposed method has a better clustering result that traditional methods, and it has some improvement on tolerance of outliers and prevention of error propagation.

  • 【会议录名称】 全国网络与信息安全技术研讨会论文集(上册)
  • 【会议名称】全国网络与信息安全技术研讨会
  • 【会议时间】2007-07
  • 【会议地点】中国山东青岛
  • 【分类号】TP393.08
  • 【主办单位】信息产业部互联网应急处理协调办公室
节点文献中: 

本文链接的文献网络图示:

本文的引文网络