节点文献

自解释神经网络的对抗鲁棒性研究

Research on the Adversarial Robustness of Self-Explaining Models

【作者】 李想

【导师】 董恺; 闫志恒;

【作者基本信息】 东南大学 , 计算机技术(专业学位), 2024, 硕士

【摘要】 神经网络近年来飞速发展,但其仍然难以在高风险决策领域中应用。一方面,神经网络容易遭受对抗攻击的影响,在模型输入上添加一个微小的扰动即可致使模型产生误分类,这促使了模型对抗鲁棒性研究的发展。另一方面,神经网络无法为其做出的决策提供合理的解释,使得神经网络缺乏透明度、灵活性和可信性,这促使了可解释人工智能研究的发展,特别是通过修改模型结构来提高模型解释能力的自解释技术。相较于另外一类不改变模型结构的后验解释技术,自解释技术通过改变模型结构来使自解释模型生成解释和预测,使得模型的对抗鲁棒性必然发生改变,因此十分有必要去探究自解释模型的对抗鲁棒性。目前还没有工作对自解释模型的对抗鲁棒性进行研究。本文系统性地对自解释技术得到的自解释模型进行对抗鲁棒性研究,具体包括以下四个方面:(1)从传统模型视角下对自解释模型对抗鲁棒性评估这一问题进行探究。在综合考虑影响模型对抗鲁棒性的六点因素之后,设计并提出一个自解释模型对抗鲁棒性评估框架以系统性评估这些修改了自身结构以获得解释能力的自解释模型在不考虑模型解释带来的影响下模型对抗鲁棒性变化。实验证明了探究自解释模型对抗鲁棒性这一问题十分复杂,但大多数情况下自解释模型相较于其骨架网络拥有更差的对抗鲁棒性。(2)从自解释模型视角下对自解释模型对抗鲁棒性评估这一问题进行探究。在分析了自解释模型为用户提供的解释可以降低传统对抗攻击的隐蔽性从而被人工检测后,设计了一种即考虑图像隐蔽性又考虑解释隐蔽性的对抗攻击名为最小解释偏移的适应性对抗攻击,并基于此攻击对自解释模型的对抗鲁棒性进行评估。实验证明了最小解释偏移的适应性对抗攻击的有效性,并得出结论:1)从对抗鲁棒性的角度来看,图像隐蔽性和解释隐蔽性之间存在冲突;2)自解释模型更容易受到针对图像隐蔽性的攻击。(3)对自解释模型对抗鲁棒性优化这一问题进行探究。本文从自解释模型结构出发,为其实现特殊的非端到端对抗训练。实验证明了端到端对抗训练对于自解释模型来说不是最优的对抗训练方法,非端到端对抗训练可以让自解释模型获得相较于用端到端对抗训练得到的对抗鲁棒性相当甚至更好的对抗鲁棒性。(4)基于上述理论研究设计了自解释模型对抗鲁棒性评估与优化原型系统。对原型系统的测试证明了本原型系统的可用性。综上所述,本文从两个视角系统分析了自解释模型的对抗鲁棒性,并为自解释模型设计了非端到端对抗训练方法以提升自解释模型对抗鲁棒性,最终设计并实现了一个自解释模型对抗鲁棒性评估与优化原型系统,为安全的可解释技术研究与发展做出贡献,推动了自解释模型的落地与应用。

【Abstract】 Neural networks have developed rapidly in recent years,yet there are still difficulties in applying them in high-stakes decision-making applications.On the one hand,neural networks are vulnerable to adversarial examples,where adding a small perturbation to the input can make the model misclassify,promoting the development of model adversarial robustness.On the other hand,neural networks fail to provide reasonable explanations for their decisions,resulting in a lack of transparency,flexibility,and trustworthiness.This leads to the recent growth in explainable artificial intelligence research,especially in self-explaining techniques that improve the interpretability of a model by modifying its architecture.In contrast to another category of explainable artificial intelligence research that does not modify model structures,self-explaining techniques necessarily modify model structures to enable self-explaining models to generate explanations.Consequently,the model’s adversarial robustness is likely to change accordingly,making it imperative to investigate the adversarial robustness of self-explaining models.Currently,there is no research examining the adversarial robustness of self-explaining models.This thesis systematically investigates the adversarial robustness of self-explaining models obtained through such techniques,focusing on four aspects:(1)Exploring the evaluation of adversarial robustness of self-explaining models from the perspective of traditional end-to-end models.After considering six factors affecting model robustness,a framework is proposed to systematically assess changes in model adversarial robustness resulting from modifications for interpretability.Experimental results reveal the complexity of this exploration,but in most cases,self-explaining models exhibit poorer adversarial robustness compared to their backbone.(2)Exploring the evaluation of adversarial robustness of self-explaining models from the perspective of self-explaining models.Analysis shows how explanations provided by these models can reduce the covertness of traditional adversarial attacks,making them detectable by humans.An adaptive adversarial attack with minimal explanation variance is designed to assess model robustness,considering both image and explanation stealthiness.Results demonstrate the effectiveness of this attack and highlight the conflict between image and explanation covertness.(3)Exploring methods for improving adversarial robustness of self-explaining models.This thesis implements specialized non-end-to-end adversarial training for self-explaining models,showing that traditional end-to-end adversarial training may not be optimal.Non-end-to-end training can achieve comparable or superior robustness,enhancing the effectiveness of self-explaining models.(4)Building upon theoretical research,a prototype system is designed to evaluate and improve the adversarial robustness of self-explaining models.Testing confirms the usability of this system.In summary,this paper systematically analyzes the adversarial robustness of self-explaining models and introduces non-end-to-end adversarial training methods to improve their robustness.Ultimately,a prototype system for evaluating and improving the adversarial robustness of self-explaining models is designed and implemented.These contribute to the advancement of secure and explainable artificial intelligence techniques,facilitating the practical application and utilization of self-explaining models.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2026年 02期
  • 【分类号】TP183
节点文献中: