节点文献

共现潜在语义向量空间模型的进一步研究

Enhanced Considerations on Co-occurrence Latent Semantic Vector Space Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 牛奉高李星

【Author】 Niu Fenggao;Li Xing;School of Mathematical Sciences,Shanxi University;Institute of Management and Decision,Shanxi University;

【机构】 山西大学数学科学学院山西大学管理与决策研究所

【摘要】 [目的/意义]文献的向量表示是文献聚类的首要任务。共现潜在语义向量空间模型(CLSVSM)通过共现分析挖掘特征词对间的最大潜在语义信息对向量空间模型(VSM)进行了语义补充,与向量空间模型相比明显提高了中文文献的聚类性能。然而,对该模型的研究还有待深入:该模型对英文文献的聚类适用性尚需检验;是否可以考虑利用除max统计量以外的其它统计量构建模型?聚类效果又会如何?面对大量的文献数据,模型的维度往往较高,运算成本大,所以有必要对模型进行优化处理。[方法/过程]首先将CLSVSM用于对英文文献集(数据来源于Web of Science,简记为WOS)的主题聚类并与VSM的聚类结果进行比较;然后利用除max统计量以外的三个常用统计量min,ave,med构建相应的CLSVSM模型,并用这四个统计量构建的CLSVSM模型对中英文文献进行聚类比较。更重要的是,我们提出了截尾共现潜在语义向量空间模型(TCLSVSM)并检验其聚类性能。[结果/结论]实验显示:CLSVSM对英文文献聚类同样适用;四种统计量构建的模型中CLSVSM-max对中英文文献的聚类效果最佳;TCLSVSM不仅能保证聚类性能,而且能显著降低运算成本。

【Abstract】 [Purpose/Significance]The vector representation of literature is a fundamental problem for efficient literature clustering. Co-occurrence latent semantic vector space model( CLSVSM) provides a semantic supplement for vector space model( VSM) by mining maximum latent semantic information between pairs of feature words,which significantly improves the clustering performance of Chinese literature in comparison to VSM.However,the study of the model still needs enhanced considerations: The applicability for the English literature clustering needs to be tested; Whether or not there are other statistics that perform better than the one in CLSVSM; With a large amount of literature data,the dimension of the model and the computational cost are often high,so it is necessary to optimize the model.[Method/Process]The paper first applied CLSVSMto topic clustering of English Literature( Web of Science,WOS,as the data source) and compared the clustering result with VSM. Then,CLSVSMmodels were constructed by using min,ave and med estimators respectively. Subsequently clustering comparison was carried out for the four models( including the CLSVSMmodel constructed using max) on both Chinese and English literature. In addition,a truncated co-occurrence latent semantic vector space model( TCLSVSM) was proposed to reduce the computational cost.[Result/Conclusion]Experiments showed that: CLSVSMis also suitable for clustering English literature; In the models constructed by the four statistics,CLSVSM-max has the best clustering performance on both Chinese and English literature; TCLSVSMcan not only guarantee the clustering performance,but also reduce the computational cost significantly.

【基金】 国家自然科学基金项目“共现潜在语义向量空间模型及其语义核的构建与应用研究”(编号:71503151);山西省高等学校创新人才支持计划“基于潜在语义的文本信息主题深度聚类研究”(编号:2016052006)的研究成果之一
  • 【文献出处】 情报杂志 ,Journal of Intelligence , 编辑部邮箱 ,2017年12期
  • 【分类号】G353.1;TP391.1
  • 【被引频次】5
  • 【下载频次】189
节点文献中: 

本文链接的文献网络图示:

本文的引文网络