节点文献
用于Web文档聚类的基于相似度的软聚类算法
A Similarity-based Soft Clustering Algorithm for Web Documents
【摘要】 提出了一种基于相似度的软聚类算法用于文本聚类,这是一种基于相似性度量的有效的软聚类算法,实验表明通过比较SISC和诸如K-means的硬聚类算法,SISC的聚类速度快、效率高。最后展望了文本挖掘在信息技术中的发展前景。
【Abstract】 This paper proposes similarity-based soft clustering (SISC), an efficient soft clustering algorithm based on a given similarity measure used in document clustering. Comparison with existing hard clustering algorithms like K-means, the experiment indicates SISC is both efficient and effective, and this algorithm is available for document clustering. In the end, it highlights the upcoming challenges of document mining and the opportunities it offers.
【关键词】 Web文本挖掘;
文本聚类;
软聚类;
相似度;
【Key words】 Web document mining; Document clustering; Soft clustering; Similarity;
【Key words】 Web document mining; Document clustering; Soft clustering; Similarity;
【基金】 教育部博士点基金项目(20030486045)“遥感影像数据库语义生成中的层次差别方法”
- 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2006年02期
- 【分类号】TP391.1
- 【被引频次】14
- 【下载频次】471