节点文献

基于大语言模型的认知推理方法研究

Research on Approaches of Cognitive Reasoning Based on Large Language Models

【作者】 刘亮

【导师】 李寿山; 张栋;

【作者基本信息】 苏州大学 , 人工智能(专业学位), 2025, 硕士

【摘要】 在人工智能领域,认知推理作为模拟人类思维过程的关键技术,它能够使机器在面对复杂问题时,像人类一样进行逻辑推理、知识联想和问题求解,这对于推动智能系统的发展具有深远的影响。然而,实际应用往往为零样本场景,即要求模型在没有特定领域标注数据的情况下,对新任务进行认知推理,这对模型的泛化和知识迁移能力是一项巨大挑战。随着大语言模型的出现,其凭借海量文本数据训练,具备丰富语言知识和一定推理能力,为解决零样本认知推理问题带来新契机。但如何充分发挥大语言模型在零样本场景下的认知推理潜力,仍是一个亟待深入研究的问题。为此,本研究旨在探索无标注样本下,结合大语言模型的认知推理方法。研究共分为三个阶段,各阶段具体内容如下:首先,现有研究往往忽略了大语言模型本身所蕴含的海量经验,而这些经验实际上能够为模型的认知推理提供有力支持。针对这一问题,本文提出了一种基于生成式思维链的认知推理方法,通过大语言模型生成若干相似样本,进而探索这些样本对模型推理的影响。具体而言,该方法在模型认知推理过程中,首先借助大语言模型自动构造若干个与用户查询相关的相似样本,然后对这些样本进行归纳总结,从而实现举一反三的效果。实验表明,该方法能够有效利用模型自身的经验,提升模型的认知推理能力。其次,针对上一阶段所提方法中单模型知识覆盖有限,难以全面掌握解决问题所需知识的问题,本文提出了一种基于多模型知识融合的认知推理方法。该方法在原有研究的基础上,通过引入了多个模型的异质知识,探究多模型间异质知识融合对认知推理的影响。具体而言,该方法首先采用两种不同的大语言模型来分别抽取与用户查询相关的知识,如语言学,背景,常识等。然后将这些知识用于具体的推理过程,并探索了两种不同的知识融合策略,即显式融合策略和隐式融合策略。实验结果表明,该方法能够有效融合多个模型的异质知识,进而提升模型对任务的理解与推理能力。最后,针对前两个研究阶段主要关注模型推理能力提升而忽视其内在机制的问题,本文从认知科学的角度出发,提出了一种基于表征工程的认知推理方法。具体而言,该方法针对两种经典的思维链方法进行剖析,分别为零样本思维链和少样本思维链,探究二者独特提示词提升模型推理能力的原理。实验结果表明,该方法不仅能进行推理错误定位,同时还能增强模型推理的稳健性,使其不受提示词影响。

【Abstract】 In the field of AI,cognitive reasoning is a key technology for simulating human think-ing processes.It enables machines to perform logical reasoning,knowledge association and problem solving like humans when facing complex problems,which has a far-reaching im-pact on promoting the development of intelligent systems.However,practical applications are often zero-shot scenarios,that is,the model is required to perform cognitive reasoning on new tasks without labeled data in a specific field,which is a huge challenge to the generaliza-tion and knowledge transfer capabilities of the model.With the emergence of large language models(LLMs),they have rich language knowledge and certain reasoning capabilities based on massive text data training,which brings new opportunities to solve the problem of zero-shot cognitive reasoning.However,how to give full play to the cognitive reasoning potential of LLMs in zero-shot scenarios is still a problem that needs to be studied in depth.To this end,this study aims to explore cognitive reasoning methods with LLMs under unlabeled samples.The research is divided into three stages,and the specific contents of each stage are as follows:Firstly,existing works often ignore the massive experience contained in the LLM itself,which can provide strong support for the cognitive reasoning of the model.To address this problem,this paper proposes a cognitive reasoning method based on a generative chain-of-thought,which generates several similar samples through an LLM,and then explores the impact of these samples on model reasoning.Specifically,in the process of model cognitive reasoning,this method first uses an LLM to construct several similar samples related to user queries automatically,and then summarizes these samples to achieve the effect of drawing inferences from one example.Experiments show that this method can effectively utilize the model’s own experience and improve the model’s cognitive reasoning ability.Secondly,in response to the problem that the single model in the method proposed in the previous work has limited knowledge coverage and it is difficult to grasp the knowledge re-quired for problem-solving fully,this paper proposes a cognitive reasoning method based on multi-model knowledge fusion.Based on the original research,this method introduces het-erogeneous knowledge of multiple models to explore the impact of heterogeneous knowledge fusion between multiple models on cognitive reasoning.Specifically,this approach first uses two different LLMs to extract knowledge related to user queries,such as linguistics,back-ground,commonsense,etc.Then this knowledge is used in the specific reasoning process,and two different knowledge fusion strategies are explored,namely explicit fusion strategy and implicit fusion strategy.The experimental results show that this method can effectively integrate the heterogeneous knowledge of multiple models,thereby improving the model’s understanding and reasoning ability of the task.Finally,in view of the problem that the first two research works mainly focused on im-proving the model’s reasoning ability while ignoring its internal mechanism,this study pro-posed a cognitive reasoning method based on representation engineering from the perspective of cognitive science.Specifically,this method analyzes two classic chain-of-thought meth-ods,namely zero-shot Co T and few-shot Co T,and explores the principle of their unique prompt text to improve the model’s reasoning ability.The experimental results show that this method can not only locate reasoning errors,but also enhance the robustness of model reasoning so that it is not affected by prompt text.

  • 【网络出版投稿人】 苏州大学
  • 【网络出版年期】2026年 07期
  • 【分类号】TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络