节点文献

中文Web检索中聚类算法的改进

Improvement of clustering algorithm in chinese web retrieval

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 耿玉良陈家琪王咏梅

【Author】 GENG Yu-liang,CHEN Jia-qi,WANG Yong-mei(College of Computer Engineering,Shanghai University of Technology,Shanghai 200093,China)

【机构】 上海理工大学计算机工程学院上海理工大学计算机工程学院 上海200093上海200093上海200093

【摘要】 对基于混合相似度的HTFC算法进行改进,要做的预处理是:建立向量空间模型,计算文档和链接的混合相似度。算法过程是:首先随机选取√kn个文档进行层次聚类,直到剩k个聚簇为止;对这k个聚簇不断迭代直到集合元素不再变化为止;然后表示出每类;最后通过用户对结果的反馈使得新生成的簇继续迭代,最终满足用户需求。算法第1步采用的是改进的k-means算法,可提高运行效率。反馈机制对原有模型进一步修正,从而提高精度。

【Abstract】 Improvement of HTFC algorithm based on mixed similarity is engaged.Pre-processes to be done are: buildingup vector space model,computing mixed similarity according to text and hyperlink.Procedure of algorithm is: firstly choose√kn texts at random,ag-glomerative clustering is executed until the number of clusters is leftk,secondly iteration is repeated until elements inthe setkeep stability;then show each class;lastly thefeedback to result can iterate again to stabilize newly cluster.By adoption of improved k-means algorithm,performance can be enhanced.The improvement of feedback to prototype can also upgrade precision.

【基金】 上海市教育委员会科研基金项目(04EB12)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2005年10期
  • 【分类号】TP391.3
  • 【被引频次】14
  • 【下载频次】241
节点文献中: 

本文链接的文献网络图示:

本文的引文网络