节点文献
基于相似度线性加权方法的检索结果聚类研究
Study on the Retrival Results Clustering Based on Linear Weighting Method of Similarity
【Author】 Liu Haibo,Zheng Dequan,Zhao Tiejun MOE-MS Key Laboratory of Natural Language Processing and Speech,Harbin Institute of Technology,Harbin 150001
【机构】 教育部-微软语言语音重点实验室哈尔滨工业大学;
【摘要】 对检索结果的聚类能够便于用户在大量搜索结果中快速找到需要的信息,传统文本聚类技术在检索结果聚类上取得的效果并不好。Lingo算法采用LSI(潜在语义索引)对检索结果进行聚类,其首先生成候选标签,然后分配文档,形成聚类。本文提出一种在Lingo算法的基础上,融合HowNet语义相似度和余弦相似度线性加权的Single-Pass改进方法对聚类进行融合和簇再发现,并提取簇标签。该方法在聚类的纯度和F值方面均取得了较好的实验结果。
【Abstract】 The retrieval results clustering can facilitate the users in finding the needed from massive information.But the effect of the traditional text clustering has been verified no good.Lingo Algorithm,which adopts LSI(Latent Semantic Indexing) for clustering,generates candidate labels first,then distributes the documents,and form the clusters finally.On the basis of Lingo Algorithm,this paper presents a linear weighted method of Single-Pass improvement,which integrates HowNet semantic similarity and cosine similarity,fuses and rediscovers clusters,and extracting the cluster labels.The experiments have showed it has good performance in purity and F-measure of clusters.
【Key words】 text clustering; information retrieval; lingo algorithm; semantic similarity; cosine similarity;
- 【会议录名称】 中国计算语言学研究前沿进展(2009-2011)
- 【会议名称】第十一届全国计算语言学学术会议
- 【会议时间】2011-08-20
- 【会议地点】中国河南洛阳
- 【分类号】TP391.3
- 【主办单位】中国中文信息学会