节点文献

基于Scopus检索和TFIDF的论文关键词自动提取方法

Keyphrases automatic extraction from the abstracts of English scientific papers based on Scopus retrieval

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈列蕾方晖

【Author】 Chen Lielei;Fang Hui;School of Electronics Science and Engineering,Nanjing University;

【机构】 南京大学电子科学与工程学院

【摘要】 客观准确的关键词能够帮助电子数据库对科研文献进行分类,也能帮助研究人员缩小文献检索的范围.提出基于TFIDF(Term Frequency-Inverse Document Frequency)与Scopus数据库检索的方法自动提取英文科研文献的关键词,将Scopus数据库包含的所有文档作为语料库,并利用Scopus API实现库内自动检索.相对于传统的人工建立并标记语料库,该方法更方便,可用数据更丰富.该方法利用摘要冗余信息量少的特点,结合全文信息的统计特征从摘要中提取关键词;考虑并建立了摘要的结构特征词,通过统计引入了短语的位置特征并进行加权,还扩展了两类停用词库用于过滤干扰词.实验结果表明该方法具有较好的性能.

【Abstract】 Keyphrases automatic extraction technology has been gradually used in scientific publications.Objective and accurate keyphrases are utilized to clustering documents in databases.In addition,suitable keyphrases assist researchers in finding relevant papers.This paper proposes a method based on TFIDF(Term Frequency-Inverse Document Frequency)and Scopus database retrieval to extract keyphrases automatically from abstracts of English scientific papers.Our method considers all the documents indexed in the Scopus database as corpus,and uses Scopus API to retrieve candidates in the database automatically.Compared with the traditional approaches that rely on manually established and annotated corpus,our method is more convenient with richer available data.Taking the advantages that abstracts have less redundant information,the key phrases were extracted from the abstracts based on the statistical features of the full text.We constructed the structural characteristics of abstracts and introduced the position feature of candidates.Moreover,two type stop-words lists for excluding noise candidates were developed for a better performance.The experimental results show that our method performed well.

  • 【文献出处】 南京大学学报(自然科学) ,Journal of Nanjing University(Natural Science) , 编辑部邮箱 ,2018年03期
  • 【分类号】TP391.1
  • 【被引频次】17
  • 【下载频次】383
节点文献中: 

本文链接的文献网络图示:

本文的引文网络