节点文献
从WOS地址字段提取二级机构数据的半自动数据清洗方法
A Semi-automatic Data Cleaning Method for Extracting Secondary Institutions’ Data from WOS Address Field
【摘要】 各高校都需要统计本校各个二级机构Web of Science(WOS)发文情况,论文提出一种基于正则表达式的半自动数据清洗方法,可从WOS地址字段中提取出发文机构排名、所属二级机构名称以及对应作者群,并以2015年南京师范大学WOS发文统计为例,进行实证研究,分析出各院系发文情况和作者发文情况。
【Abstract】 Chinese higher education institutions need to count the articles included in Web of Science(WOS) by their secondary institutions. This paper puts forward a semi-automatic data cleaning method based on regular expressions for extracting ranking of the dispatch agency, name of the secondary institutions and the corresponding authors from WOS address fields. At last, it takes the statistics of articles included in WOS of Nanjing Normal University in 2015 as an example to conduct an empirical study, and analyze the situation of the articles issued by various faculties and authors.
【Key words】 Secondary institutions; Regular expression; Data cleaning; WOS address field; Sci-tech novelty search;
- 【文献出处】 新世纪图书馆 ,New Century Library , 编辑部邮箱 ,2017年08期
- 【分类号】G353.1
- 【被引频次】4
- 【下载频次】264