节点文献

一种分布式爬虫系统的设计与应用

Design and Application of a Distributed Crawler System

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 周逸李新陈远平

【Author】 Zhou Yi;Li Xin;Chen Yuanping;Computer Network Information Center, Chinese Academy of Sciences;University of Chinese Academy of Sciences;

【机构】 中国科学院计算机网络信息中心中国科学院大学

【摘要】 文献计量学是一种把握学科发展态势的定量分析方法。传统基于文献计量学的研究步骤需手动操作且流程繁琐,针对这一问题,设计并实现了一种基于scrapy-redis分布式爬虫的学科发展态势分析系统。该系统包含了1.负责爬取并解析web of science文献数据的数据预处理层。解决了由于网速不稳定造成的爬虫丢失网页问题,保障数据完整性。设计了一种动态计算参考文献所属学科分布情况的算法2.基于Django搭建的结果展示层,通过web服务向用户展示学科态势分析结果。用户只需输入初始待爬取页面的URL即可通过web服务获得学科态势分析结果。该系统为文献计量学提供了一种更便捷、更快速、扩展性高的分析手段。

【Abstract】 Bibliometrics is a quantitative analysis method to master the trend of discipline development. In traditional bibliometrics research, the research steps are manual and tedious. In order to improve it, an analysis system based on scrapy-redis, a distributed crawler framework, are designed and implemented to master the trend of discipline development. The system includes: 1.data processing layer.It is responsible for crawling literature information on web of science. The problem that spiders lose some web pages due to unstable network speed is solved,ensuring data integrity. Design an dynamic algorithm to calculate the discipline distribution according to references.2. web service layer. It is made by Django and used for showing the analysis result. User only needs to input the initial URL that user wants to crawl,then the user just need to wait for the analysis result and require it by web service. The system provides a more convenient, more rapid and high scalability approach to do bibliometrics research.

  • 【文献出处】 科研信息化技术与应用 ,e-Science Technology & Application , 编辑部邮箱 ,2019年01期
  • 【分类号】G350;TP391.3;TP311.13
  • 【下载频次】118
节点文献中: 

本文链接的文献网络图示:

本文的引文网络