节点文献

基于分类和关键词组抽取的信息检索算法

An Information-retrieval Algorithm Based on Classification and Key Phrase Extraction

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 钟敏娟; 林亚平; 陈治平;

【Author】 ZHONG Min-juan,LIN Ya-ping,CHEN Zhi-ping(College of Computer and Communication,Hunan University,Changsha Hunan 410082,China)

【机构】 湖南大学计算机与通信学院; 湖南大学计算机与通信学院 长沙410082; 长沙410082; 长沙410082;

【摘要】 本文提出一种基于分类和关键词组抽取的信息检索算法。该算法利用文本分类和信息抽取技术辅助检索,避免了向量空间模型算法中时间复杂度过大,查准率不高的缺点。针对传统的信息检索性能指标无法有效地衡量检索结果的排序状况,本文还引入了排序误差率概念用于评价检索结果的排序。实验结果表明,所提算法与TFIDF算法、基于分类的交互式检索算法相比,具有更快的查询速度,更高的查准率和更小的排序误差率。

【Abstract】 In this paper, a new information retrieval algorithm based on classification and key phrase extraction is proposed. Compared with traditional vector space model, this algorithm reduces time complexity and improves precision using of text classification and information extraction. Then a new performance criterion named ranking error is contributed to solve the problem that the traditional performance evaluation methodology cant evaluate the ranking results of retrieved documents efficiently. The experiment result shows that the proposed algorithm outperforms TF*IDF and Interactive Retrieval based on classification in speed, precision and ranking error.

【基金】 国家自然科学基金(60272051)
  • 【文献出处】 系统仿真学报 ,Acta Simulata Systematica Sinica , 编辑部邮箱 ,2004年05期
  • 【分类号】TP391.3
  • 【被引频次】34
  • 【下载频次】452
节点文献中: 

本文链接的文献网络图示:

本文的引文网络