节点文献
基于标点符号插入的中文对抗文本生成方法
Chinese adversarial text generation method based on punctuation insertion
【摘要】 自然语言处理模型容易受到对抗文本的影响。现有的针对中文对抗文本的生成方法主要采用形近字替换或同音字替换等方式,在面对鲁棒性较强的预训练模型时,对抗文本中的扰动数增加导致文本的流畅性和可读性下降,从而生成的对抗文本质量差。而英文对抗文本中符号插入方法不能完全适用于中文。同时,在黑盒场景下,先验知识的缺乏导致难以生成高质量的对抗文本。面向中文文本分类任务,提出了一种基于标点符号插入的中文对抗文本生成方法。在黑盒设置中,该方法采用一种词性重要性计算方法,并结合标点符号插入设计了一种适用于中文的字符级扰动方法,以实现对抗文本的生成。实验证明,针对文本分类任务,所提方法在使用两个真实数据集训练的LSTM(long short-term memory)和BERT(bidirectional encoder representations from transformers)模型上攻击成功率显著提升,同时,成功避免了直接破坏原始句子,保持了原始句意。在测试中,该方法能够达到97%的语义相似度,明显优于基线方法。
【Abstract】 The susceptibility of natural language processing models to adversarial texts has been a significant concern. Current methods for generating adversarial texts in Chinese were mainly based on replacing characters with visually similar or homophonic ones. However, when faced with robust pre-trained models, these methods led to increased perturbations in adversarial texts, resulting in reduced fluency and readability, and thus generating lowquality adversarial texts. Moreover, symbol insertion methods used in English adversarial texts were not entirely applicable to Chinese. Additionally, in a black-box scenario, the lack of prior knowledge made it difficult to generate high-quality adversarial texts. A punctuation-based method for generating adversarial texts for Chinese text classification tasks was proposed. Under a black-box setting, a novel part-of-speech importance calculation was utilized and combined with punctuation insertion to design a character-level perturbation approach suitable for Chinese, achieving the generation of adversarial texts. Experiments were conducted, and the results demonstrated that for text classification tasks, the proposed method significantly improved the attack success rate on LSTM and BERT models trained with two real-world datasets. Furthermore, the method successfully avoided direct destruction of the original sentences and maintained the original meaning. In the tests, a semantic similarity of up to 97% was achieved, which was significantly better than the baseline methods.
【Key words】 Chinese text classification; adversarial text generation; black-box attack;
- 【文献出处】 网络与信息安全学报 ,Chinese Journal of Network and Information Security , 编辑部邮箱 ,2025年02期
- 【分类号】TP391.1
- 【下载频次】4