节点文献
基于语义模型的文档特征提取
Feature Extraction of Document Based on Semantic Model
【摘要】 文档特征提取是文本检索领域研究的最重要的问题之一。论文提出了一种全新的文档特征表示方式—语义模型。使用WordNet分析语义,提取主题句向量组用以确保文档含义的准确表达,再综合成文本向量保证特征表示的相关性。采用这种方式对文档作特征提取能在一定程度上同时提高文本检索的查全、查准率。理论分析与实验结果均表明论文的基于语义模型的文档特征提取方法是可行且有效的。
【Abstract】 Feature extraction is one of the most important issues in text retrieval.This paper presents a new way to denote the feature of document-semantic model.It uses WordNet to analyse the meaning of words:hypernym,hyponym,antonym and synonym,draw topic sentence vectors to obtain the precise expression of the meaning of documents,then integrate them into text vector to ensure the relativity of the denotation of feature.The method based on semantic model can improve the rate of precision and recall of IR meanwhile.Both theory and experiment results show that the method of feature extraction of documents based on semantic model is feasible and efficient.
【Key words】 text retrieval; feature extraction; semantic model; WordNet;
- 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2005年17期
- 【分类号】TP391.1
- 【被引频次】8
- 【下载频次】339