节点文献
基于python的互联网数据爬取与解析的研究与实践
Research and Practice of Internet Data Crawling and Parsing Based on Python
【摘要】 随着信息的多元化和大数据时代的到来,人们在生活中对网络的应用越来越广泛,使得网络拥有了海量的数据。如何在庞大的网络数据中高效快速地获取对用户有用的信息是一项尤为重要的技术。笔者着重研究了网络数据爬取技术中基于Python语言第三方库的网络爬虫技术,并尝试利用该技术对部分网站数据进行爬取、解析和重新建构。
【Abstract】 With the diversification of information and the arrival of the era of big data, people are more and more widely used in life, making the network has a large amount of data. How to efficiently and quickly acquire useful information for users in huge network data is a particularly important technology. The author focuses on the research of web crawler technology based on Python third-party library, and tries to use this technology to crawl, parse and reconstruct some web site data.
【关键词】 python第三方库;
数据爬取;
数据解析;
数据保存;
【Key words】 python third-party library; Data crawling; Data analysis; Data preservation;
【Key words】 python third-party library; Data crawling; Data analysis; Data preservation;
【基金】 吉林省大学生创新创业训练计划项目“互联网数据爬取与解析”(项目编号:201810202108)
- 【文献出处】 信息与电脑(理论版) ,China Computer & Communication , 编辑部邮箱 ,2019年17期
- 【分类号】TP393.092;TP391.3
- 【被引频次】3
- 【下载频次】472