节点文献

基于知识蒸馏的深度神经网络模型不确定性校准方法研究

Research on Uncertainty Calibration Methods for Deep Neural Networks Based on Knowledge Distillation

【作者】 杨洋;

【导师】 周学海; 王超;

【作者基本信息】 中国科学技术大学 , 计算机科学与技术, 2025, 博士

【摘要】 不确定性校准作为深度学习领域内一个新兴的研究方向,其核心目标在于赋能深度神经网络模型,使其能够精确量化预测结果所伴随的不确定性。近年来,随着深度神经网络模型的持续演进与精进,这些模型已逐渐渗透至高风险决策领域,成为指导关键行动的重要工具。以自动驾驶为例,一辆集成了深度神经网络决策系统的自动驾驶车辆,通过摄像头捕捉的图像识别技术,能够精准地检测行人及各类障碍物。当决策模型面临难以确切判断障碍物存在与否的情境,即预测结果的不确定性显著升高时,车辆需更加倚重备用传感器的数据输入或引入人工监控,以确保在制动等关键操作上做出准确无误的决策。在此情境下,深度神经网络模型不仅要承担基本预测任务,还需准确量化预测结果的不确定性,为整个系统的综合判断提供必要支持。然而,深度神经网络模型固有的训练机制常致使其输出的不确定性量化值与实际偏差显著。鉴于此,开发高效且实用的不确定性校准方法,对于促进深度神经网络模型在高风险领域的稳健应用具有至关重要的作用。当前研究动向揭示,基于知识蒸馏的不确定性校准方法因其能够在保持模型结构不变的同时增强预测精确度和不确定性评估能力,而在深度神经网络领域受到广泛关注。因此,本文致力于运用知识蒸馏方法探究深度神经网络模型的不确定性校准策略,并在此基础上提出了一系列解决方案。本文的主要工作与贡献可概括如下三个方面:首先,针对图像分类与目标检测两大核心任务,本文提出了一种融合了集成蒸馏方法与对比表征学习的新型不确定性校准框架。此框架引入了反例行为作为一类独特的知识范畴,从而在知识传递的维度上实现了对传统知识蒸馏方法的显著拓展。具体而言,本文提出的基于对比的双向学习方法,不仅高度重视并致力于学习集成教师模型在量化模型输出不确定性方面的卓越能力,而且能够识别并纠正学生模型中不准确的不确定性量化表现,有效规避未来类似错误的重演。实验结果表明,本文所提出的融合了集成蒸馏方法与对比表征学习的不确定性校准框架,在不确定性校准性能方面展现出了相较于其他现有方法的显著优越性。其次,针对集成蒸馏方法在目标检测任务中知识传递效率低下问题,本文提出了一种专为目标检测任务设计的、基于知识概率化的集成蒸馏框架。通过一系列的探究实验,本文证明了在目标检测场景下,相较于传统的单纯数值表示方法,知识概率化方法在不确定性表征的精确度、模型预测性能的增强以及知识传递的鲁棒性等多个维度上均展现出了显著的优势。鉴于此,本文在新构建的集成蒸馏框架中融入了知识概率化的知识传递机制。该机制通过对目标检测任务中的教师模型与学生模型的多样化输出数值执行概率化处理,将所得的概率分布特征融入训练损失函数中,调控整个集成蒸馏的训练流程。实验数据显示,采用基于知识概率化的集成蒸馏框架,能够显著增强神经网络模型在目标检测任务中的不确定性量化表现。最后,为了应对集成蒸馏方法因训练成本高昂而难以在资源受限环境中广泛应用的挑战,本文提出了一种基于自蒸馏的高效不确定性校准框架。该框架摒弃了传统的教师-学生训练架构,转而采取自我蒸馏的方式,即利用模型自身产生的知识作为训练的参考基准,从而实现了训练资源的大幅节省。经过一系列详尽的实验验证,本文所提出的基于自蒸馏的高效不确定性校准框架,在多个关键性能指标上显著优于那些具有相似训练成本的不确定性校准框架。综上所述,本论文以知识蒸馏方法为基础,深入探索并成功挖掘出深度神经网络模型在不确定性校准方面的解决方案。通过这一系列精心设计的框架和方法,本论文为深度神经网络模型在安全关键领域的应用与发展作出了贡献。

【Abstract】 Uncertainty calibration,as an emerging research direction in the field of deep learning,aims to empower deep neural network models with the ability to accurately quantify the uncertainty associated with their predictions.In recent years,with the continuous evolution and refinement of deep neural network models,these models have increasingly permeated high-stakes decision-making domains,becoming essential tools for guiding critical actions.For instance,in autonomous driving,a vehicle equipped with a deep neural network-based decision system can utilize image recognition technologies to accurately detect pedestrians and various obstacles.When the decision model encounters scenarios where the presence of obstacles is ambiguous—leading to heightened uncertainty in its predictions—the vehicle must rely more heavily on auxiliary sensor inputs or introduce human oversight to ensure accurate and reliable decisions for critical actions such as braking.In such cases,the deep neural network model is tasked not only with basic prediction but also with accurately quantifying the uncertainty of its predictions,thereby providing necessary support for the system’s overall judgment.However,the inherent training mechanisms of deep neural network models often lead to a significant discrepancy between their quantified uncertainty outputs and the actual uncertainty.Consequently,the development of efficient and practical uncertainty calibration methods is critical for enabling the robust application of deep neural network models in high-risk domains.Recent research trends highlight that uncertainty calibration methods based on knowledge distillation have garnered significant attention in the field of deep neural networks.These methods enhance prediction accuracy and uncertainty evaluation capabilities while maintaining the original model architecture.Motivated by these advancements,this dissertation focuses on exploring uncertainty calibration strategies for deep neural network models using knowledge distillation.Building on this foundation,it proposes a series of innovative solutions.The main contributions of this work can be summarized in the following three aspects:First,for the core tasks of image classification and object detection,this dissertation introduces an innovative uncertainty calibration framework that integrates ensemble distillation techniques with contrastive representation learning.This framework pioneers the incorporation of counterexample behaviors as a unique category of knowledge,significantly expanding the traditional dimensions of knowledge transfer in conventional distillation methods.Specifically,the proposed contrastive bidirectional learning approach not only emphasizes learning the ensemble teacher models’ superior capability in quantifying uncertainty but also identifies and corrects inaccurate uncertainty quantification in the student models.This correction effectively prevents similar errors from recurring.Experimental results demonstrate that the proposed framework,which combines distillation techniques and contrastive representation learning,exhibits superior performance in uncertainty calibration compared to existing methods.Second,addressing the inefficiency of knowledge transfer in ensemble distillation methods for object detection tasks,this dissertation proposes a novel probabilistic knowledge-based ensemble distillation framework specifically designed for object detection.Through a series of exploratory experiments,this study demonstrates that,in the context of object detection,the probabilistic representation of knowledge outperforms traditional numerical representation methods in multiple dimensions,including the accuracy of uncertainty representation,enhanced prediction performance,and the robustness of knowledge transfer.Based on these findings,this dissertation integrates a probabilistic knowledge transfer mechanism into the newly constructed ensemble distillation framework.This mechanism probabilistically processes the diverse outputs from teacher and student models in object detection tasks,incorporating the resulting probability distribution features into the training loss function to regulate the entire ensemble distillation training process.Experimental data show that the proposed probabilistic knowledge-based ensemble distillation framework significantly enhances uncertainty quantification performance in neural network models for object detection tasks.Finally,to address the high training costs of ensemble distillation methods,which limit their widespread application in resource-constrained environments,this dissertation proposes an innovative self-distillation-based efficient uncertainty calibration framework.This framework abandons the traditional teacher-student training paradigm and instead adopts a self-distillation approach,using knowledge generated by the model itself as the reference for training.This approach significantly reduces training resource requirements.Comprehensive experimental validation demonstrates that the proposed self-distillation-based framework achieves superior performance across multiple key metrics compared to uncertainty calibration frameworks with similar training costs.In summary,this dissertation,grounded in knowledge distillation techniques,explores and develops innovative solutions for uncertainty calibration in deep neural network models.Through the meticulously designed frameworks and methods presented,this work contributes to the application and advancement of deep neural network models in safety-critical domains.

  • 【分类号】TP183
节点文献中: 

本文链接的文献网络图示:

本文的引文网络