节点文献

基于因果推理的知识图谱偏差去除研究

Causal Inference-Based Debiasing Research on Knowledge Graph

【作者】 任林

【导师】 欧阳纯萍;

【作者基本信息】 南华大学 , 电子信息(专业学位), 2024, 硕士

【摘要】 在知识图谱构建的过程中,模型的性能在很大程度上取决于训练数据的质量和多样性。但知识图谱相关的数据集的收集过程繁琐,而且数据的标注往往需要消耗大量的人力时间资源。此外,这些数据集大多由自然语言文本构成,其中固有的偏差难以进行量化分析,因此难以主动去除数据集中存在的偏差。为了应对这一问题,本文利用因果推理技术,针对知识图谱构建过程中的实体关系联合抽取任务中的数据不平衡和选择性偏差提出了OVO框架(Covariance and Variance Optimization Framework);针对知识图谱补全任务中由模型过度拟合知识图谱结构导致的偏差问题提出了CIDF框架(Causal Inference-Based Debiasing Framework)。具体研究内容如下:(1)对于实体关系联合抽取任务,OVO针对数据不平衡和选择性偏差对模型的准确性和泛化能力的影响,通过最小化相同标签样本的方差以及特征之间的协方差,来优化模型的特征表示,使得模型能够更公平地学习到各类别之间的差异,从而提高在样本数量较少的类别上的性能。本文将OVO应用到基线模型Sp ERT.PL、PURE-F和PL-Marker中,来评估其在实体关系联合抽取任务中的有效性。针对评估指标Rel+,OVO在数据集Sci ERC上相较于原本基线提升了2.2%、3.3%和2.9%,在数据集ACE2005上相较于原本基线提升了3.2%、0.3%和1.2%。(2)对于知识图谱补全任务,CIDF关注由模型过度拟合知识图谱结构导致的偏差问题。这种过度拟合会降低模型在预测缺失的信息时的总体性能。为了缓解这一问题,CIDF将其细分为了深度偏差和广度偏差,然后通过因果推理中的干预和反事实方法对模型的训练目标和推理方式进行优化,旨在尽可能保证精确度的前提下提高模型的整体性能并减少偏差的影响。对于评价指标MRR,将CIDF应用到基线模型Sim KGC中,在数据集WN18RR、FB15k-237、Wikidata5M-Trans和Wikidata5M-Ind提升了1.6%、1.0%、2.9%和1.2%。综上所述,本文针对实体关系联合抽取和知识图谱补全任务进行了深入研究,通过使用因果推理技术进行偏差分析及偏差缓解方案设计,为知识图谱中的偏差缓解提供了新的思路。

【Abstract】 In the field of knowledge graphs,the performance of models largely depends on the quality and diversity of the training data.However,the process of collecting datasets related to knowledge graphs is cumbersome,and annotating the data often requires significant amounts of time and computational resources.Moreover,these datasets are mostly composed of natural language texts,in which inherent biases are difficult to quantify for analysis,thus making it challenging to proactively eliminate biases present in the datasets.In order to address this problem,this paper proposes the OVO for the data imbalance and selective bias in the task of joint entity-relationship extraction in the process of knowledge graph construction by using causal inference techniques;and proposes the CIDF for the bias problem in the task of knowledge graph supplementation caused by the model’s overfitting to the structure of the knowledge graph.The specific research contents are as follows:(1)For the joint entity and relation extraction,this paper thoroughly investigates the impacts of data imbalance and selection bias on the model’s performance and generalization ability.To mitigate the issues brought by these biases,this paper optimizes the model’s feature representation by minimizing the variance of samples with the same label and the covariance among features.This approach enables the model to learn the differences between categories more equitably,thereby enhancing its performance on categories with fewer samples.By applying OVO to the baseline models Sp ERT.PL,PURE-F and PLMarker.For the assessment metric Rel+,OVO improved by 2.2%,3.3% and2.9% on the dataset Sci ERC and by 3.2%,0.3% and 1.2% on the dataset ACE2005,compared to the original baseline.(2)For the knowledge graph completion,this paper focuses on the bias problem caused by models overfitting the structure of knowledge graphs.Such overfitting can diminish the overall performance of the model in predicting missing information.To mitigate this issue,the paper subdivides it into in-depth bias and in-breadth bias.It then optimizes the model’s training objectives and inference methods through interventions and counterfactual approaches in causal inference,aiming to enhance the model’s overall performance and mitigate the impact of bias,all while ensuring accuracy as much as possible.For the evaluation metric MRR,applying CIDF to the baseline model Sim KGC boosted 1.6%,1.0%,2.9%,and 1.2% in datasets WN18 RR,FB15k-237,Wikidata5M-Trans,and Wikidata5M-Ind.In summary,this paper provides an in-depth study for the task of joint entity and relation extraction and knowledge graph completion,and provides new ideas for bias mitigation in knowledge graphs.

  • 【网络出版投稿人】 南华大学
  • 【网络出版年期】2025年 07期
  • 【分类号】TP391.1
节点文献中: 

本文链接的文献网络图示:

本文的引文网络