节点文献
生物信息学工具知识图谱构建和新实体嵌入生成
Knowledge Graph Construction for Bioinformatics Tools and Generation of New Entity Embeddings
【作者】 李华;
【导师】 时小虎;
【作者基本信息】 吉林大学 , 软件工程, 2022, 硕士
【摘要】 生物信息学是当今生命科学和自然科学的前沿领域,其主要研究内容之一为开发或设计一系列相关工具,以更有效地获取、分析和管理各种生物数据,为相关科研人员提供便捷的数据信息渠道。近年来,随着该领域的快速发展,生物信息学领域的优秀工具不断涌现。与此同时,生物信息学细分领域众多,产生的相关工具种类繁杂,对学习和使用工具造成了一定难度。知识图谱可以帮助人们解决上述问题。谷歌率先提出知识图谱的相关概念,用以辅助数据的存储、分析、决策等,现已广泛应用于不同领域。然而,在生物信息学领域,尚未有针对生物信息学工具的知识图谱出现。通过构建生物信息学工具知识图谱,能够沉淀更多的专业领域知识,帮助搜索和推荐,以及辅助实现更精准的问答系统,具有很强的实用价值。为了进行知识推理和知识挖掘等下游任务,需要先将知识图谱表示成嵌入表示。生物信息学相关工具软件的开发速度很快,意味着所构建的知识图谱需要持续迭代更新,新的实体会不断涌现,因此如何在下游任务中对新实体进行有效表示是知识图谱应用的难点之一。本文利用知识图谱的相关技术,设计和构建了生物信息学工具知识图谱,并针对出现的新实体提出了一种新实体嵌入式生成表示方法NEEGAT,论文的具体工作包括:(1)知识图谱构建及可视化针对生物信息学工具构建领域知识图谱。首先,使用selenium和Scrapy等自动化技术和爬虫技术,获取了工具、作者、工具所属领域、论文、期刊、关键词、引用等信息,进行对齐、筛选、清洗、去重和降噪。其次,将数据进行拆解,形成三元组。最后,引入图数据库,实现知识图谱的可视化,最终形成了一个拥有近四万实体、二十万三元组的庞大的知识图谱。(2)提出了一种基于图注意力网络的新实体嵌入表示算法针对动态更新的知识图谱不断涌现的新实体,为了避免重新训练整个知识图谱,提出了一种基于图注意力网络的新实体嵌入表示算法NEEGAT。算法使用Trans E进行预训练,获取工具图谱的三元组的整体语义信息,利用逻辑注意力将知识图谱以外部知识的方式引入,使用多头图注意力网络进一步整合邻居节点间多种维度的链接关系。此外,基于所构建的知识图谱构建了Bio Tools数据集,并针对链接预测和三元组分类两种下游任务进行采样生成实验数据集,以对本文方法进行检验。实验结果表明,在链接预测任务和三元组分类任务上,本文提出的NEEGAT方法和对比方法相比,在整体上均取得了最佳表现,说明了该算法能更好地解决知识图谱的新实体嵌入生成问题。
【Abstract】 Bioinformatics is the frontier field of life sciences and natural sciences today,one of the main contents of which is to develop and design a series of relevant tools to make various biological data effective to acquire,analyze and manage,providing relevant researchers with convenient data information access ways.With the rapid development of this field in recent years,excellent tools of Bioinformatics have been emerging.At the same time,due to the numerous sub-fields of Bioinformatics,the variety of related tools is complicated,which makes it difficult for people to use and learn.Knowledge graph is expected to solve the problems above.Since Google put forward the relevant concept,knowledge graph has been widely used in different fields to assist data storage,data analysis,decision-making and so on.However,in the field of bioinformatics,no knowledge graph for Bioinformatics tools has emerged yet.Knowledge graph of Bioinformatics tools has a strong practical value,as it can precipitate more professional domain knowledge,and can help search,recommend,and assist to achieve more accurate Q&A system.In order to carry out downstream tasks such as knowledge graph reasoning and knowledge mining,it is necessary to use embeddings to represent knowledge graph.The rapid development of Bioinformaticsrelated tools and software means that the constructed knowledge graph need to be continuously iterated and updated,and new entities will emerge frequently.Therefore,how to effectively represent new entities in the downstream tasks is one of the difficulties in the application of knowledge graph.This paper designs and constructs the Bioinformatics tool knowledge graph by using the related technology,and proposes a new entity embedding generation method NEEGAT for the emerged new entities.The main works of this paper are as follow:(1)Knowledge graph construction and visualizationBuild a domain knowledge graph for Bioinformatics tools.First,using automated technologies and crawler technologies such as selenium and Scrapy,the tools,authors,fields to which the tools belong,papers,journals,keywords,citations,and other information are obtained,aligned,screened,cleaned,deduplicated and denoised.Second,the data is disassembled to form a triple.Finally,a graph database is introduced to visualize the knowledge graph,and finally a huge knowledge graph with nearly40,000 entities and 200,000 triples was formed.(2)Development of new entity embedding generation algorithmConsidering new entities that keep emerging from the dynamically updated knowledge graph,in order to avoid retraining the whole knowledge graph,a new entity embedding representation algorithm NEEGAT based on graph attention network is proposed.The algorithm uses Trans E for pre-training to obtain the overall semantic information of the triples of the tool graph,uses logical attention to introduce the knowledge graph in the form of external knowledge,and uses the multi-head graph attention network to further integrate the multi-dimensional link relationships between neighbor nodes.In addition,a Bio Tools dataset is constructed based on the constructed knowledge graph and an experimental dataset is generated by sampling the two downstream tasks of link prediction and triplet classification to test the method in this paper.The experimental results show that on the link prediction task and triple classification task,the NEEGAT method proposed in this paper has achieved the best overall performance compared with the comparison method,which shows that the algorithm can better solve the new entity embedding generation problem of knowledge graph.
【Key words】 Bioinformatics; Knowledge Graph; Knowledge Graph Embedding; Graph Neural Network; Attention Mechanism;
- 【网络出版投稿人】 吉林大学 【网络出版年期】2023年 01期
- 【分类号】TP391.1;Q811.4