节点文献

面向小样本的知识图谱自动构建方法研究

Research on Few-Shot Knowledge Graph Automatic Construction Methods

【作者】 高超

【导师】 武剑洁;

【作者基本信息】 华中科技大学 , 软件工程(专业学位), 2024, 硕士

【摘要】 在接触网施工领域,技术交底文档对于确保施工的质量和安全至关重要。传统的技术交底文档多依赖人工编写,导致文档质量参差不齐。基于知识图谱自动创建技术交底文档,可以有效提高文档编写质量和效率。然而,接触网领域的知识图谱生成受到专业性强、样本稀缺等因素影响,面临信息抽取困难、模型泛化能力差等问题,为此展开小样本的知识图谱自动构建方法研究。针对专业领域信息抽取困难的问题,提出基于多任务学习的联合抽取模型,通过共享底层特征信息来提高信息抽取的准确性。模型在传统字符嵌入的基础上融入词级信息,以丰富字符的语义表达。捕获上下文信息结合句子特征,构建共享特征向量。提出基于正负样本的主题实体识别方法,增强实体抽取的精确性。结合语义特征矩阵,完成客体和谓词的抽取,形成完整的信息三元组。针对小样本关系抽取模型泛化能力差、稀有关系抽取困难的问题,进一步提升联合抽取模型在小样本情况下的性能,提出基于角度间隔策略的关系抽取模型,通过对于损失函数的优化,增强模型对于不同关系的识别。模型融合迁移学习和领域适应策略,以增强特征抽取的效率和准确性。在传统的交叉熵损失函数基础上引入角度间隔策略,优化特征空间中关系类别的区分度,增强模型对不同关系类别的区分能力。在公开数据集WebNLG和自建数据集OTD上展开对比实验,关系分类结果的F1值分别达到93.7%和83.7%,表明模型能够有效识别接触网领域中的实体与关系;在公开数据集Few Rel上,5-way-1-shot和10-way-1-shot的准确率比性能第二的模型分别提升0.52%和0.48%,展示了模型在小样本情况下仍能有效学习并泛化到未见过关系的能力。

【Abstract】 In the field of overhead contact system(OCS)construction,technical disclosure documents are crucial for ensuring the quality and safety of construction.Traditionally,these documents are manually written,which leads to uneven document quality.Automatically generating technical disclosure documents based on knowledge graphs can effectively improve the quality and efficiency of document writing.However,the generation of knowledge graphs in the OCS field is affected by factors such as high specialization and sample scarcity,resulting in difficulties in information extraction and poor model generalization ability.For this reason,research on few-shot knowledge graph automatic construction methods is carried out.A joint extraction model based on multi-task learning is proposed to improve the accuracy of information extraction by sharing underlying feature information and solve the problem of difficult information extraction in specialized fields.The model incorporates word-level information on the basis of traditional character embeddings to enrich the semantic expression of characters.A shared feature vector is constructed by capturing contextual information combined with sentence features.A topic entity recognition method based on positive and negative samples is proposed to enhance the precision of entity extraction.Combining the semantic feature matrix,the extraction of subjects,predicates and objects is completed,forming complete information triplets.A relation extraction model based on angular margin strategy is proposed to enhance the model’s recognition of different relationships and the performance of joint extraction models under the condition of small samples by optimizing the loss function,which solves the poor generalization ability and difficulty in extracting rare relationships of few-shot relation extraction models.The model integrates transfer learning and domain adaptation strategies to enhance the efficiency and accuracy of feature extraction.Introducing an angular margin strategy on the basis of the traditional cross-entropy loss function optimizes the discriminability of relationship categories in the feature space,enhancing the model’s ability to distinguish between different relationship categories.Comparative experiments are carried out on the public dataset WebNLG and the selfbuilt dataset OTD.The F1 scores for relation classification reach 93.7% and 83.7%,respectively,which indicates that the model’s effective recognition of entities and relationships in the OCS field.On the Few Rel dataset,the accuracies of 5-way-1-shot and10-way-1-shot are higher than that of the model with the second highest performance by0.52% and 0.48%,demonstrating the model’s ability to effectively learn and generalize to unseen relationships even in few-shot scenarios.

  • 【分类号】U225;TP391.1;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络