节点文献

基于跨语言广义向量空间模型的跨语言文档聚类方法

Cross-Lingual Document Clustering Based on Similarity Space Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 唐国瑜夏云庆张民郑方

【Author】 TANG Guoyu1,XIA Yunqing1,ZHANG Min2,ZHENG Thomas Fang 1(1.Department of Computer Science and Technology,Tsinghua University,Beijing 100084,China; 2.Institute for Infocomm Research,A-STAR,138632,Singapore)

【机构】 清华大学计算机科学与技术系资讯通信研究院

【摘要】 跨语言文档聚类主要是将跨语言文档按照内容或者话题组织为不同的类簇。该文通过采用跨语言词相似度计算将单语广义向量空间模型(Generalized Vector Space Model,GVSM)拓展到跨语言文档表示中,即跨语言广义空间向量模型(Cross-Lingual Generalized Vector Space Model,CLGVSM),并且比较了不同相似度在文档聚类下的性能。同时提出了适用于GVSM的特征选择算法。实验证明,采用SOCPMI词汇相似度度量算法构造GVSM时,跨语言文档聚类的性能优于LSA。

【Abstract】 Cross-Lingual Document Clustering is the task to automatically organize a large collection of cross-lingual documents into groups according to their contents or topics.This work extends traditional monolingual Generalized Vector Space Model(GVSM) to Cross-Lingual GVSM(CLGVSM) by using cross-lingual term similarity calculation methods in order to represent documents in different languages and compare different term similarity calculation methods in cross-lingual document clustering.This work also proposes new feature selection method for CLGVSM.Experiment results show that GVSM with Second Order Co-occurrence Point wise Mutual Information(SOCPMI) term similarity measure outperforms the latent semantic analysis(LSA) method.

【基金】 科技部资助项目(2009DFA12970)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2012年02期
  • 【分类号】TP391.1
  • 【被引频次】11
  • 【下载频次】269
节点文献中: 

本文链接的文献网络图示:

本文的引文网络