节点文献
作物高光谱领域文献知识图谱构建关键技术研究
Research on Key Technologies for Constructing Knowledge Graph in Crop Hyperspectral Literature
【作者】 刘霞;
【导师】 曹晓兰;
【作者基本信息】 湖南农业大学 , 农业硕士(专业学位), 2024, 硕士
【摘要】 农业是国民经济的基石,也是人类生存和发展的根本。随着大数据技术在农业生产中的深入应用,以及农业知识数据库的持续扩充,面对农业知识的查询与利用,研究者们往往陷入效率困境,导致此类知识的实际应用潜力未能有效释放。鉴于知识图谱在知识组织与实践应用中的显著优势,近年来,一些学者已着手构建服务于农业领域的知识图谱,且已取得若干实质性进展。然而,目前在作物高光谱领域的知识图谱研究,仍显不足。故此,本研究致力于整合该领域内零散的作物高光谱知识文献资料,借助知识图谱的结构性表述机制,以期对作物高光谱实体间的多元关联实现有效管理,从而有针对性地化解当前所面临的知识分布离散、共享受阻及关联梳理不充分等问题。具体研究内容如下:(1)作物高光谱本体与数据集构建本研究采取自上而下的路径来构建作物高光谱本体。首先,本研究参照了相关的既有数据集,并携手行业专家对作物高光谱知识体系进行了深度剖析。在此基础上,利用Protégé软件构建了作物高光谱领域的本体模型,确定了8类实体、7种关系,及实体和关系的属性。通过构建本体,明晰了知识提取的边界条件,为构建作物高光谱专业数据集提供了基础。继而,融合专家见解及已成型的本体架构,对数据进行了分类与标注,构建了作物高光谱数据集。(2)作物高光谱实体关系联合抽取模型构建构建了基于注意力与门控机制,融入词汇嵌入信息的作物高光谱领域实体关系联合抽取模型(A Joint Extraction Model of Entity Relationships in Crop Hyperspectral Domain Based on Attention and Gating Mechanisms for Embedding Vocabulary Information,CH_AGM)。为克服传统串行抽取模型存在的误导入词信息及三元组重叠弊端,该模型在运用BERT提取字符层面特征之余,进一步融入词汇层面信息以丰富字符特征表达。同时,借助注意力机制,将各类关系信息精准地整合至句子表征之中。此外,运用门控机制来调控关系信息的融入强度,以确保所融入信息的适宜度与有效性。在解码阶段,结合了潜在关系预测来进一步提升模型的预测能力。实验结果表明:在自建作物高光谱数据集上,模型的P、R、F1值达到了80.5%、81.6%和82.4%。相比于流行联合抽取模型Cas Rel和PRGC在P、R、F1上分别提升了0.4%、2.2%、0.8%和1.8%、0.9%、2.9%。(3)作物高光谱知识图谱构建选择Neo4j图形数据库作为存储载体,用于高效保存经抽取与融合处理后的实体三元组数据。以此存储框架为基础,设计开发了一个面向作物高光谱领域的知识图谱应用系统,并初步实现了基础查询功能,为用户提供多样化的知识服务。
【Abstract】 Agriculture is the cornerstone of the national economy and the fundamental basis for human survival and development.With the in-depth application of big data technology in agricultural production and the continuous expansion of agricultural knowledge databases,researchers often face efficiency difficulties in querying and utilizing agricultural knowledge,resulting in the failure to effectively unleash the practical application potential of such knowledge.Given the significant advantages of knowledge graphs in knowledge organization and practical applications,in recent years,some scholars have begun to construct knowledge graphs that serve the agricultural field,and have made several substantial progress.However,there is still insufficient research on knowledge graphs in the field of crop hyperspectral analysis.Therefore,this study aims to integrate scattered literature on crop hyperspectral knowledge in the field,utilizing the structural representation mechanism of knowledge graphs,in order to effectively manage the diverse relationships between crop hyperspectral entities,and thus address the current problems of knowledge distribution dispersion,hindered sharing,and insufficient association sorting in a targeted manner.The specific research content is as follows:(1)Crop hyperspectral ontology and dataset constructionThis study adopts a top-down approach to construct crop hyperspectral ontology.Firstly,this study referred to relevant existing datasets and conducted in-depth analysis of the crop hyperspectral knowledge system in collaboration with industry experts.On this basis,an ontology model in the field of crop hyperspectral was constructed using Protégésoftware,and 8 types of entities,7 relationships,and the attributes of entities and relationships were determined.By constructing the ontology,the boundary conditions for knowledge extraction were clarified,providing a foundation for constructing crop hyperspectral professional datasets.Subsequently,by integrating expert insights and established ontology architecture,the data was classified and annotated,and a crop hyperspectral dataset was constructed.(2)Construction of a joint extraction model for crop hyperspectral entity relationshipsA Joint Extraction Model of Entity Relationships in Crop Hyperspectral Domain Based on Attention and Gating Mechanisms for Embedded Vocabulary Information was proposed.To solve the problem of introducing incorrect word information and entity overlap in traditional pipeline extraction models,the model further embeds vocabulary information on the basis of extracting character feature information using BERT to enrich character features.Simultaneously,through the integration of attention mechanisms,various relational information types are efficiently incorporated into sentence representations.Furthermore,a gating system is implemented to modulate the extent of relational information assimilation,ensuring the rationality and effectiveness of the information.In the decoding stage,directional prediction is combined to further enhance the predictive ability of the model.The experimental results show that on the self built crop hyperspectral dataset,the P,R,and F1values of the model reached 80.5%,81.6%,and 82.4%.Compared to the popular joint extraction models Cas Rel and PRGC,they improved P,R,and F1 by 0.4%,2.2%,0.8%,and1.8%,0.9%,and 2.9%,respectively.(3)Fabrication of a Crop Hyperspectral Knowledge RepositoryChoose the Neo4j graphic database as the storage medium for efficiently storing entity triplet data after extraction and fusion processing.Based on this storage framework,a knowledge graph application system for crop hyperspectral fields was designed and developed,and basic query functions were initially implemented to provide users with diverse knowledge services.
【Key words】 Knowledge Graph; Crop Hyperspectral; Joint Extraction; Deep Learning; Neo4j;
- 【网络出版投稿人】 湖南农业大学 【网络出版年期】2025年 11期
- 【分类号】TP391.1;S126