节点文献
基于词间语义相关度的搜索结果聚类算法
An Algorithm for Search Result Clustering Based on Semantic Relevance Between Words
【摘要】 将查询结果根据内容进行聚类是提高搜索引擎服务质量的关键技术之一.搜索结果聚类时只能从文档标题和文档片段中抽取有限信息,传统聚类方法难以准确计算其相似度.提出了一种基于词间语义相关度的搜索结果聚类算法,该算法以词为聚类的核心,词所出现的文档为词的属性,根据词在搜索结果文档中共现的情况来划分类别.该方法可以充分利用词间的语义相关性,类别划分后即可确定类名.实验结果表明,对搜索结果聚类时与K-Means和STC算法相比,质量上有所提高.
【Abstract】 To automatically group search results into thematic clusters has become an important topic in the search engine research area,which can be used to improve the quality of searching service.Since the search engine only returns a ranked list of documents along with their partial content(snippets),simply porting the well known generic algorithms does not work well,because the amount of data for the clustering algorithm is often extremely small and low-quality.An algorithm for search result clustering based on semantic relevance between words is proposed.Word is the clustering elements rather than snippets,and its attributes are the snippets where it appears.To fully explore the semantic relevance between words,the snippets are clustered according to the graph of the semantic relevance.The cluster name can be given once the cluster algorithm is finished.The experiment shows that the algorithm performs better than K-Means and STC.
【Key words】 search result clustering; semantic relevance between words; document similarity;
- 【文献出处】 郑州大学学报(理学版) ,Journal of Zhengzhou University(Natural Science Edition) , 编辑部邮箱 ,2009年01期
- 【分类号】TP391.1;TP18
- 【被引频次】2
- 【下载频次】200