节点文献
基于跨语言广义向量空间模型的跨语言文档聚类方法
Cross-lingual Document Clustering Based on Similarity Space Model
【Author】 Tang Guoyu~1,Xia Yunqing~1,Zhang Min~2,Thomas Fang Zheng~1 1 Department of Computer Science and Technology,Tsinghua University,Beijing 100084 2 Institute for Infocomm Research,A-STAR,Singapore
【机构】 清华大学计算机科学与技术系; 资讯通信研究院;
【摘要】 跨语言文档聚类主要是将跨语言文档按照内容或者话题组织为不同的类簇。本文通过采用跨语言词相似度计算将单语广义向量空间模型(Generalized Vector Space Model,GVSM)拓展到跨语言文档表示中,即跨语言广义空间向量模型(CLGVSM),并且比较了不同相似度的在文档聚类下的性能。同时提出了适用于GVSM的特征选择算法。实验证明,采用SOCPMI词汇相似度度量算法构造GVSM时,跨语言文档聚类的性能优于LSA。
【Abstract】 Cross-lingual Document Clustering is the task to automatically organize a large collection of cross-lingual documents into groups according to their contents or topics.This work extends traditional monolingual Generalized Vector Space Model(GVSM) to Cross-lingual GVSM(CLGVSM) by using cross-lingual term similarity calculation methods in order to represent documents in different languages and compare different term similarity calculation methods in cross-lingual document clustering.This work also proposes new feature selection method for CLGVSM.Experiment results show that GVSM with Second Order Co-occurrence Pointwise Mutual Information(SOCPMI) term similarity measure outperforms the latent semantic analysis(LSA) method.
【Key words】 cross-lingual document clustering; text similarity; document clustering;
- 【会议录名称】 中国计算语言学研究前沿进展(2009-2011)
- 【会议名称】第十一届全国计算语言学学术会议
- 【会议时间】2011-08-20
- 【会议地点】中国河南洛阳
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会