节点文献

基于语义模型的文档特征提取

Feature Extraction of Document Based on Semantic Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李开荣林颖杭月芹

【Author】 Li Kairong Lin Ying Hang Yueqin (College of Information Engineering,Yangzhou University,Yangzhou 225009)

【机构】 扬州大学信息工程学院计算机科学与工程系扬州大学信息工程学院计算机科学与工程系 扬州225009扬州225009扬州225009

【摘要】 文档特征提取是文本检索领域研究的最重要的问题之一。论文提出了一种全新的文档特征表示方式—语义模型。使用WordNet分析语义,提取主题句向量组用以确保文档含义的准确表达,再综合成文本向量保证特征表示的相关性。采用这种方式对文档作特征提取能在一定程度上同时提高文本检索的查全、查准率。理论分析与实验结果均表明论文的基于语义模型的文档特征提取方法是可行且有效的。

【Abstract】 Feature extraction is one of the most important issues in text retrieval.This paper presents a new way to denote the feature of document-semantic model.It uses WordNet to analyse the meaning of words:hypernym,hyponym,antonym and synonym,draw topic sentence vectors to obtain the precise expression of the meaning of documents,then integrate them into text vector to ensure the relativity of the denotation of feature.The method based on semantic model can improve the rate of precision and recall of IR meanwhile.Both theory and experiment results show that the method of feature extraction of documents based on semantic model is feasible and efficient.

【关键词】 文本检索特征提取语义模型WordNet
【Key words】 text retrievalfeature extractionsemantic modelWordNet
【基金】 江苏省高校自然科学基金项目(编号:02KJB520013)
  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2005年17期
  • 【分类号】TP391.1
  • 【被引频次】8
  • 【下载频次】339
节点文献中: 

本文链接的文献网络图示:

本文的引文网络