节点文献

生成式文本信息隐藏和隐写分析技术研究

Research on Information Hiding and Steganalysis Technology of Generative Text

【作者】 钟山;

【导师】 任延珍;

【作者基本信息】 武汉大学 , 网络空间安全, 2021, 硕士

【摘要】 基于文本的信息隐藏技术是信息隐藏领域中一个重要的研究方向。相比图像、音视频等载体,文本在公开信道的传输过程中基本不会受到重压缩操作的影响,具备更强的多信道适应能力。随着文本生成技术的快速发展,基于生成的文本隐写算法成为新的研究热点。相比基于修改的文本隐写算法,生成式文本隐写算法具备更高的嵌入容量,可广泛应用于隐蔽通信、隐蔽存储等场景。同时,为了防止该类技术被恶意滥用,需要对相关检测技术展开研究,以用于非法信息的拦截和破坏。而且,该类检测技术还可以作为隐写算法的安全性评估标准。因此,针对生成式文本信息隐藏和隐写分析技术的研究具有重要的意义。本文一方面深入研究隐写文本生成过程,设计整体性能更好的文本隐写算法;另一方面,通过语义特征提取和分析,提出检测效果更好的文本隐写分析算法。本文的主要研究内容和创新点如下:1)针对载密文本和自然文本分布差异过大的问题,本文提出一种基于自适应Top-k采样的文本隐写算法。结合语言生成模型输出的条件概率分布信息,以Top-k截断前后的分布差异为参照,在不同生成时刻自适应地选择合适的k值,从而更好地保持原始分布。实验结果证明,本文方案相比现有固定k值的采样方式提升了隐写算法的整体性能表现。2)考虑到生成式文本隐写算法受到文本体裁的影响,本文探索分析了不同类型隐写文本的隐写特性,并为隐蔽通信应用中载体形式的研究提供了参考方案。结合模型迁移的思想,本文将几种最新文本隐写算法应用于古诗词、对联和文言文三种中文体裁形式中,并对比分析了相同隐写算法下各类隐写文本的性能差异。实验结果表明,即使在相同嵌入方式下,不同类型隐写文本的性能表现也大有不同。3)针对现有隐写分析方法无法有效利用文本上下文信息的问题,本文提出一种基于预训练语言模型的文本隐写分析算法。以GPT-2、BERT等预训练模型作为特征提取器,再将提取得到的结合上下文信息的隐层特征作为下游文本隐写分析网络的输入,从而更好地捕捉到文本序列的特征差异,再结合注意力机制,完成可疑样本的检测。实验结果表明,相比现有隐写分析方法,本文方案具有更好的检测性能。

【Abstract】 Text-based information hiding technology is an important research topic in the field of information hiding.Compared with images,audio,video and other carriers,text is basically not affected by recompression operation in the transmission process of public channel,and thus has stronger multi-channel adaptability.With the rapid development of text generation technology,text steganography algorithm based on generation has become a new research hotspot.Compared with the modified text steganography algorithm,the generative text steganography algorithm has higher embedding capacity,and can be widely used in covert communication,covert storage and other scenarios.At the same time,in order to prevent this kind of technology from being maliciously abused,it is necessary to carry out research on relevant detection technology to intercept and destroy illegal information.Moreover,this kind of detection technique can also be used as a security evaluation standard for steganographic algorithms.Therefore,the research on generative text information hiding and steganographic analysis is of great significance.On the one hand,this paper deeply studies the process of steganographic text generation and designs a text steganography algorithm with better overall performance.On the other hand,through semantic feature extraction and analysis,a better detection effect of text steganography analysis algorithm is proposed.The main research contents and innovations of this paper are as follows:1)Aiming at the problem that the distribution of ciphertext and natural text is too different,this paper proposes a text steganography algorithm based on adaptive Top-K sampling.Combined with the conditional probability distribution information output by the language generation model,and taking the distribution difference before and after top-k truncation as reference,the appropriate K value is selected adaptively at different generation moments,so as to better maintain the original distribution.Experimental results show that the proposed scheme improves the overall performance of the steganography algorithm compared with the existing sampling method with fixed K value.2)Considering that the generative text steganography algorithm is influenced by the text genre,this paper explores and analyzes the steganography characteristics of different types of steganographic texts,and provides a reference scheme for the study of carrier forms in covert communication applications.Combined with the idea of model transfer,this paper applies several latest steganographic algorithms to three Chinese genres: ancient poems,couplets and classical Chinese,and compares and analyzes the performance differences of various steganographic texts under the same steganographic algorithms.The experimental results show that the performance of different types of steganographic text is very different even under the same embedding mode.3)Aiming at the problem that the existing steganalysis methods cannot effectively use text context information,this paper proposes a text steganalysis algorithm based on a pre-trained language model.Pre-trained models such as GPT-2 and BERT are used as feature extractors,and then the hidden layer features combined with context information are extracted as the input of the downstream text steganalysis network,so as to better capture the feature differences of the text sequence.Combined with the attention mechanism,the detection of suspicious samples is completed.The experimental results show that compared with the existing steganalysis methods,the proposed scheme has better detection performance.

  • 【网络出版投稿人】 武汉大学
  • 【网络出版年期】2025年 01期
节点文献中: