节点文献
多模态大模型安全研究进展
Recent progress of the security research for multimodal large models
【摘要】 多模态大模型的安全性研究已成为当下人工智能领域的焦点。由于大模型以深度神经网络为核心构建,因此与深度神经网络类似,存在多种安全风险。此外,由于其特有的复杂性,以及广泛的应用场景,也使得大模型面临一些独特的安全风险。本文系统地总结多模态大模型的安全风险,包括对抗攻击、越狱攻击、后门攻击、版权窃取、幻觉现象、泛化问题以及偏见问题等。具体来说,在对抗攻击中,攻击者通过构造微小但具有欺骗性的对抗样本,使大模型在面对带噪输入时产生严重的误判;越狱攻击利用大模型的复杂结构,绕过或破坏原有的安全约束和防御措施,使模型执行未授权的操作,甚至泄露敏感数据;后门攻击则通过在大模型的训练阶段植入隐秘的触发器,使模型在特定条件下做出攻击者预期的反应;未经授权的窃取者可能未经模型拥有者的同意随意分发或进行商业使用,将导致模型版权拥有者遭受损失;幻觉现象,即模型输出与输入不一致的问题;泛化问题即大模型当前应对部分新数据分布或风格的能力仍显不足;大模型在性别、种族、肤色和年龄等敏感问题上的偏向性可能引发伦理等问题。随后,针对这些安全风险分别介绍相应的解决方案。本文旨在为理解和应对多模态大模型的独特安全挑战提供一个独特的视角,促进多模态大模型安全技术的发展,引导未来相关安全技术的发展方向。
【Abstract】 With the rise of deep learning technology, artificial intelligence technology has progressed from shallow machine learning to deep learning and from small-scale data learning to big data learning. In recent years, with the continuous improvement of data and computing resources, the scale of deep learning models has continued to increase, and large-scale pretrained models(large models) have begun to emerge. In general, any model that is trained on large-scale extensive data(usually via large-scale self-supervised training) and can be adjusted(e. g., through fine-tuning) to a wide range of downstream tasks and can be referred to as a large model. Meanwhile, a typical large model that is trained by utilizing multimodal information, such as text, images, videos, and audio, and aims to complete diverse multimodal application tasks is known as a multimodal large model. In the past two years, domestic and foreign multimodal large models, such as ChatGPT, Llama, and Qwen, have achieved remarkable success. As multimodal large models continuously evolve, the research on the security of these models has become the focus in the field of artificial intelligence. Given the superior processing abilities of these models in various multimodal tasks, their security issues may have harmful consequences. Large models are built with deep neural networks as their core, so they encounter security risks similar to those faced by deep neural networks. In addition, given the unique complexity of large models and their wide range of application, they face unique security risks. In this study, we systematically summarize the security risks associated with multimodal large models, including adversarial attacks, jailbreak attacks, backdoor attacks, copyright theft, hallucination phenomena, generalization issues, and bias problems. In adversarial attacks, attackers construct small yet deceptive adversarial examples to cause misjudgments by large models when they are fed with these adversarial perturbed inputs. Jailbreak attacks exploit the complex structure of large models to bypass or destroy the original security constraints and defense measures, enabling the models to perform unauthorized operations and/or even leak sensitive data in the outputs. Backdoor attacks involve implanting hidden triggers during the training phase of large models, causing the models to exhibit attacker-intended behaviors under specific conditions. Meanwhile, unauthorized thieves may distribute or use large models for commercial purposes without the consent of the model owners, causing losses to the copyright owner of the models. The hallucination phenomenon refers to the issue of inconsistency between a large model’s output and input. The generalization problem indicates the inability of large models to deal with new data distributions or styles. The bias of large models on sensitive issues, such as gender, race, skin color, and age, may lead to ethical problems, which may further produce severe consequences.After the presentations of these security risks, we introduce corresponding solutions to. By presenting the progress of research on the security risks and corresponding solutions of multimodal large models, this study aims to provide a unique perspective for understanding and addressing the unique security challenges of multimodal large models, promote the development of security technologies for multimodal large models, and guide the future direction of related security technology development. In conclusion, multimodal large models demonstrate excellent performance in many tasks and applications and provide various types of assistance to people’s work and daily lives. Through this research on the security technologies of multimodal large models, we hope to ensure the safety and reliability of these models, thereby providing a guarantee for people’s normal work and life.
【Key words】 multimodal large model; large model security; adversarial example(AE); jailbreak attack; backdoor attack; copyright theft; model hallucination; model bias;
- 【文献出处】 中国图象图形学报 ,Journal of Image and Graphics , 编辑部邮箱 ,2025年06期
- 【分类号】TP309;TP18
- 【下载频次】100