节点文献
面向中文长文本分类场景的对抗样本攻击技术
Adversarial Attack for Chinese Long Text Classification Scenarios
【摘要】 深度学习模型在多种应用中展现了卓越能力,但其面对对抗样本的脆弱性不容忽视。该文针对现有方法在扰动位置选择时不考虑扰动连续替换后的影响,且在长文本场景下需要大量查询受害模型的问题,提出了一种适用于中文长文本的对抗样本攻击方法DQNAttack。该方法设计了针对中文特性的候选扰动生成策略,并提出使用深度强化学习模型高效定位长文本样本的扰动替换位置。实验表明,该方法较现有方法在新闻领域数据集和法律领域数据集上攻击成功率提升10%左右,在语义相似度、语句流畅度和攻击效率上明显优于现有方法。
【Abstract】 Deep learning models have demonstrated excellent capabilities in various applications, but their vulnerability to adversarial samples cannot be ignored. In response to the problem that existing methods do not consider the impact of continuous replacement of perturbations when selecting perturbation positions, and require a large number of queries on the victim model in long text scenarios, an adversarial sample attack method DQNAttack suitable for Chinese long texts has been proposed. This method designs a candidate perturbation generation strategy for Chinese characteristics and proposes the use of deep reinforcement learning models to locate perturbation replacement positions. The experiment shows that the method proposed in this paper increases the success rate by about 10% compared with existing methods on news and legal datasets.
【Key words】 adversarial attack; deep reinforcement learning; mask language model;
- 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2026年02期
- 【分类号】TP391.1;TP18
- 【下载频次】9