节点文献

基于Web的异构学术信息抽取与聚合方法研究

Research on Heterogeneous Academic Information Extraction and Aggregation Based on Web

【作者】 刘子玉

【导师】 黄晶;

【作者基本信息】 吉林大学 , 计算机软件与理论, 2019, 硕士

【摘要】 互联网时代,海量网页信息层出不穷,科技学术领域更是如此。每年有大量的学术期刊论文发表,也有很多学术人物信息在互联网上公开。如果想了解某个学术期刊或学术人物,并不能轻松获得,需要在互联网上点击一系列超链接才有可能找到。对于科研人员而言,能否快速获得学术信息非常必要。在此背景下,本文研究了基于Web的异构学术信息抽取与聚合方法,提出自动化的算法框架以帮助研究人员从互联网大量的异构网页中迅速挖掘所需信息。本文的主要工作如下:1.针对基于web的学术期刊信息抽取与聚合问题,本文提出了C-HMM算法框架。该框架中的正文提取算法(Content Extraction)可提取网页中的主要信息,实现了降噪的效果;隐马尔可夫模型(HMM)可同时对多个网站进行抽取,相较于现有的启发式算法提升了模型的泛化能力。C-HMM算法框架分为三个步骤:首先,通过爬虫爬取期刊主页;然后,对主页信息进行预处理和正文提取;最后,利用HMM对期刊信息进行抽取与聚合。2.针对基于web的学术人物信息抽取与聚合问题,本文提出了F-HMM算法框架。该框架中的fastText算法可对网页信息块进行预标注,此算法解决了关键字词典无法对人物多种信息块预标注的问题;隐马尔可夫模型(HMM)刻画了信息块的时序信息,提升了模型效果。F-HMM算法框架与C-HMM框架有以下三点不同:(1)采用SVM对学术人物主页进行选择,取代期刊主页选择时采用的关键词匹配策略;(2)由于学术人物主页结构复杂,正文提取算法可能会过滤有用信息,因此舍弃;(3)采用fastText算法取代了原有的关键词匹配方法,对信息块进行预标注。3.上述两个工作是吉林省重点科技研发项目“大数据和移动互联时代的快速知识共享系统研究、开发与应用”的重要组成部分。作者将上述工作以及论文、新闻和征稿信息的自动化爬虫系统加入到《学术头条》APP的开发中,方便了研究人员快速获取学术信息。目前APP拥有7000多名用户、400多万篇论文、6000多种期刊以及670多万个学术人物,实际测试结果表明,本文工作取得了良好的效果。

【Abstract】 In the Internet era,massive web page information emerges endlessly,especially in the field of science and technology.A large number of academic journals publish papers every year,and lots of researchers publish information on the Internet at the same time.If someone wants to get information about academic journals or researchers,it is not easy to make it.He or she needs to click on a series of hyperlinks on the Internet.Meanwhile,It is necessary for researchers to get academic information quickly.In this background,this paper studies the extraction and aggregation of heterogeneous academic information on the Internet,and proposes an automated algorithm framework to help researchers quickly mine the required information from a large number of heterogeneous web pages on the Internet.This paper mainly does the following works:1.Aiming at the problem of information extraction and aggregation of web-based academic journals.In this paper,C-HMM algorithm framework is proposed.The content-extraction algorithm in the framework achieves the effect of noise reduction.And the HMM can extract multiple websites simultaneously.These improve the generalization ability of the model compared to existing heuristic algorithms.The C-HMM algorithm framework is divided into three steps.Firstly,the crawler technology is used to get the home page of academic journals.Then,after the step of data preprocessing,the content-extraction algorithm is used to extract the important information in the homepage.Finally,the HMM model is used to extract and aggregate the information of academic journals.2.Aiming at the problem of information extraction and aggregation of web-based academic researchers,this paper proposes a F-HMM algorithm framework.The fastText algorithm in the framework pre-labels the webpage information block,which solves the problem that the keyword dictionary cannot pre-label multiple information blocks of the character.Based on the fastText algorithm,the HMM model is used to describe the timing information of the information blocks,which improves the effect of the model.The F-HMM algorithm framework is different from the C-HMM framework in the three aspects: First,SVM is adopted to select the homepage of academic researchers,and replace the keywordmatching strategy.Second,due to the complex structure of the academic researchers’ homepage,the content extraction algorithm may filter useful information,so this algorithm is abandoned.Third,using the fastText algorithm can replace the original keyword matching method to pre-label the information blocks.3.The above two tasks are an important part of the research,development and application of the rapid knowledge sharing system in the era of big data and mobile internet in Jilin Province.The above work,the automatic crawler system of papers,news and CFP information are added into the development of the "Academic Headline" APP,which also facilitates researchers to quickly get academic information.At present,APP has more than7,000 users,4 million papers,6,000 journals and 6.7 million academic researchers,the effects show that the work of this paper has achieved good results.

  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2019年 10期
  • 【分类号】TP391.1;TP393.092
  • 【被引频次】1
  • 【下载频次】79
  • 攻读期成果
节点文献中: