节点文献
基于Corpus库的词语相似度计算方法
Measurement of word similarity based on Corpus
【摘要】 构建了一个语义关联库,称为Corpus库,该库使用词语空间和关系空间结构化地存储了词语和其上下文之间的统计信息,并通过阅读大量的预料数据来训练其相关数据。详细介绍了Corpus库的训练方法,并对训练过程中出现的大量关系提出了裁剪方案。在此基础上,通过构建词语的上下文关系向量提出了一种词语相似度算法。实验证明这是一种有效的对词语相似度进行计算的方法。
【Abstract】 A semantic relevant database named Corpus was built to store the required information in word similarity measurement. Corpus got the information from large scale text training and store the information in word space and relation space after analysis and tailoring. The word similarity measurement algorithm by constructing the context relation vectors based on Corpus was given, which proved to be a feasible method by experiments.
【基金】 交大数字家电实验室“Advanced information retrieval technology using the knowledge base”项目
- 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2006年03期
- 【分类号】TP391.1
- 【被引频次】67
- 【下载频次】519