节点文献

基于知识图谱的自然语言问题生成研究

Research on Natural Language Question Generation over Knowledge Graphs

【作者】 毕胜;

【导师】 漆桂林;

【作者基本信息】 东南大学 , 软件工程, 2023, 博士

【摘要】 知识图谱作为一种重要的知识表示方式,在近年来受到越来越多的关注。知识图谱通过将知识表示为图的形式,使得知识的组织、查询和利用变得更加便捷。随着知识图谱技术的发展,人们开始关注如何让计算机能够更好地理解和应对人类的自然语言查询,通过提问的方式挖掘知识图谱中的信息,并为用户提供更为个性化的知识服务。基于知识图谱的问题生成旨在根据给定的子图,即一组相连的三元组,生成可被回答的自然语言问题,被广泛应用于多种人工智能系统,如问答、对话以及在线教育等。基于规则和模板的传统方法人力成本高、可泛化能力差。随着算力的提升和可用数据的增加,基于深度学习的方法受到了研究人员的青睐。通常,自动问题生成被当作序列预测任务,采用编码器-解码器框架进行建模,相比传统方法取得了令人瞩目的进展。然而,现有方法在问题的准确性、泛化性和可控性上仍有诸多不足亟待解决。首先,三元组作为高度凝练的知识载体,其信息量不足以生成表达清晰、语法正确的自然语言问题;其次,在生成复杂问题时,模型不具备组合泛化能力,难以厘清多跳三元组之间的内在关联;最后,在生成难度可控的问题时,缺乏有效的复杂度评估方法,同时,由于已有工作对难度的建模方式单一,无法达到令人满意的可控性。基于上述讨论,本文致力于改善知识图谱问题生成三个方面的局限性:在输入信息有限的情况下,如何保证生成问题的正确性;提升模型的组合泛化能力,强化生成问题与多跳路径之间的映射关系;构建自动化难度估计算法,缓解难度标签建模单一导致的生成结果多样性低、可控性差。具体而言,本文开展了如下三个方面的研究:1.面向简单问题,提出了基于知识增强和语法引导的问题生成方法。针对输入信息量低,采用外部知识对三元组中的实体和关系进行扩充,并设计了全局关系编码器帮助理解答案实体以缓解语义漂移现象。在解码端,对每一个词进行类型预测,并将类型分布融入到当前生成的过程中。此外,通过掩码预测和自动编码器分别训练了句法树和语义依存评估器。在解码过程中,评估器将已生成序列的句法和语义依存信息用于当前生成步骤;解码完成后,对生成结果进行评价获得强化学习奖励,以缓解训练和测试阶段评测差异产生的暴露偏置。实验证明,外部知识可以让生成结果表达更清晰、多样,类型约束和强化学习帮助缓解了语义漂移和语法错误的现象。在公开数据集上的性能全面优于基线方法。2.面向复杂问题,提出了基于模块化对偶学习的问题生成方法。该方法利用知识图谱问答和问题生成之间的内部关联,设计了形式统一的对偶学习框架。首先,对于给定输入,模型通过一个离散隐变量产生一个共享网络的组织方式,即动态路由。其次,共享网络中包含多个结构相同的神经网络层,可以根据路由被放在模块化结构中的任意位置,且可以重复使用。为了更有效地利用对偶任务之间的共享机制,实现了基于损失变化量的参数传递准则,判断当前模块是否接受对偶任务共享的归纳偏置。实验证明,对偶学习框架可以同时提升复杂问答和问题生成的效果。此外,模块化共享网络可以显著增强模型的组合泛化能力,并将学到的知识进行迁移。3.面向难度可控的问题,提出了基于软模板和反事实推理的问题生成方法,以解决现有方法对难度建模单一以及复杂度标签与生成结果之间缺乏稳定的因果关系。首先,设计了基于预训练语言模型的软模板构建方法,软模板是一组可学习的参数,无需人工注释。同时,引入一个离散的动态软模板选择器,最大限度提升不同类型问题的模板多样性。然后,实现了一个子图表征解耦模块,将输入的三元组分离成与当前问题相关和无关的部分,降低了噪声对于生成结果的干扰。基于此,设计了基于反事实推理的问题生成,将软模板和分离后的事实表征结合,以连续提示学习的方式优化解码器。反事实推理能够探索修改特定属性形成的反事实样本和真实样本之间的差异,学习不同难度的提问模式。对于缺乏复杂度标注的限制,提出了基于实体、问句和问答过程的可解释复杂度估计方法。实验证明,本方法相比基线方法取得了显著领先,特别是在生成结果的难度可控能力和不同难度问题之间的多样性获得大幅提升。

【Abstract】 Knowledge graphs(KG),as an important way of representing knowledge,have received increas-ing concerns in recent years.By representing knowledge in a graph,KG makes it more convenient to organize,query,and exploit knowledge.With the advancement of KG technology,people are be-ginning to focus on enabling computers to better understand and respond to natural language queries,mine information from KG through questioning and provide more personalized knowledge services.KG-based question generation(KGQG)aims to generate answerable natural language questions based on a given subgraph,i.e.,a set of connected triples.It is widely used in various artificial intelligence systems,such as question answering(QA),dialogue,and online education.Traditional methods based on rules and templates are labor-intensive and poorly generalization capabilities.With the dramatic increase in available data and computing power,deep learning-based approaches have gained favor among researchers.Typically,automatic question generation is per-formed as a sequence prediction task and modeled using an encoder-decoder framework,which has made impressive progress compared to traditional methods.However,many things could still be improved in terms of correctness,generalization,and controllability of the questions generated by existing methodologies.First,triples,as highly condensed knowledge carriers,are not informative enough to generate well-expressed and grammatically correct natural language questions.Second,when generating complex questions,the model needs to gain the ability of compositional generaliza-tion,making it hard to clarify the intrinsic relationships between multiple triples.Finally,there is no practical complexity estimator when generating questions with controllable difficulty.Meanwhile,satisfactory controllability could not be achieved due to the simple difficulty modeling in the existing works.Motivated by the above discussions,this paper is devoted to alleviating the three limitations in KGQG: how to guarantee the correctness of generated questions with limited input information?improving the compositional generalization of the model and strengthening the mapping relationship between generated questions and multi-hop facts? building an automatic difficulty estimation and relieving the low diversity of generated results caused by an identical difficulty modeling and poor controllability.Specifically,this paper conducts research in the following three areas.1.For simple question generation,a model based on knowledge-enhanced and grammar-guided is proposed.External knowledge is employed to augment the entities and relations in the triples for limited input information.A global relation encoder is designed to facilitate understanding the answer to alleviate the semantic drift phenomenon.Each word is given a type prediction at the decoding stage,and the type distribution is incorporated into the current word gener-ation.Additionally,syntactic tree and semantic dependency evaluators are trained by mask prediction and autoencoders,respectively.Moreover,the evaluators utilize the generated se- quence’s syntactic and semantic dependency information to guide the current generation step.After decoding,the generated questions are evaluated to obtain reinforcement learning(RL)rewards,mitigating the exposure bias caused by teacher forcing during training.Experiments demonstrate that external knowledge can make the generated results more explicit and diverse.The type constraints and RL help improve semantic and syntactic correctness.The proposed model’s performance on public datasets outperforms the baselines comprehensively.2.For complex question generation,a modular dual learning-based model was proposed.The model leverages the intrinsic connection between KGQA and KGQG,and designs a uniform dual learning framework.First,for a given subgraph,the model generates a shared layer organi-zation via a discrete hidden variable,i.e.,dynamic routing.Second,the network bank contains multiple structurally identical neural network layers that can be placed anywhere in the modu-lar structure based on routing and can be reused.To effectively utilize the shared mechanism between dual tasks,a parameter transmission criterion based on the loss ratio is implemented to determine whether the current module accepts the inductive bias shared by the peer task.Exper-iments illustrate that the dual learning framework can simultaneously improve the performance of KGQA and KGQG.In addition,the modular shared network can significantly enhance the model’s compositional generalization and transfer the learned knowledge.3.For difficulty-controllable question generation,a method based on the soft template(ST)and counterfactual reasoning is proposed to solve the identical difficulty modeling and the lack of stable causal effect between complexity labels and generated results.First,an ST construction method based on a pre-trained language model is designed.The ST is a group of learnable parameters that do not require manual annotation.In parallel,a discrete dynamic ST selec-tor is introduced to maximize the diversity of templates for various question types.Then,a subgraph representation disentanglement module is implemented to decouple the input triples into relevant and irrelevant parts of the current question,reducing noise interference on the generation results.Consequently,a counterfactual reasoning module is designed,which com-bines ST and disentangled fact representations to optimize the decoder in a continuous prompt learning style.Counterfactual reasoning can explore the differences between counterfactual samples shaped by modifying specific attributes and actual samples to learn asking patterns with varying difficulty.An explainable complexity estimation method is proposed for datasets without available complexity annotation,considering the quantification of entities,interroga-tives,and QA.Experimental results prove that the proposed model significantly outperforms the baselines,especially in the controllability and diversity of the generated questions.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2025年 03期
  • 【分类号】TP391.1;TP18
节点文献中: