节点文献

基于多规则下自然语言文本隐写的研究

Research on Natural Language Text Steganography Based on Multiple Rules

【作者】 吴宁;

【导师】 李廉;

【作者基本信息】 兰州大学 , 计算机科学与技术·计算机应用技术, 2020, 博士

【摘要】 随着计算机科学及网络技术为主导的信息科学及产业的飞速发展,诸如窃听、重放、抵赖、信息泄露、非法使用等信息安全问题也日益凸显。相较于传统的信息加密技术,信息隐写技术可以以一种更加隐秘的方式保护重要的信息。信息隐写通常采用冗余量较大的载体完成,例如图像,音频,视频等。文本作为历史最悠久、使用量最大的媒体信息,由于冗余量较小而难以搭载太多秘密信息,但也因此更有研究价值。在本文中,我们以自然语言文本作为隐写研究对象,从不同的角度出发,对于自然语言文本隐写方法加以研究,具体研究内容和取得的工作成果如下:(1)开展了基于变语法规则的文本信息隐写方法研究。在研究中我们发现,不同维度的语法模型在构建过程中对于训练文本所抓取的特征不同,因此影响到最后生成的隐写文本质量。因此,我们将不同维度语法模型加以结合,以利用不同维度语法模型的各自优势。在研究中,我们提出了基于此思想的、两种不同混合语法规则的隐写模型,即基于半频交叉规则的文本信息隐写方法和基于多规则语言模型交替的文本信息隐写方法。这两种方法都是尝试在更为灵活的语言关系基础之上实现隐写文本的生成,以期待取得更好的隐写效果。我们对所提出方法完成了建模并选用了不同类别的数据进行实验,验证了方法的有效性。(2)开展了基于转移概率的文本信息隐写方法研究。在研究中我们突出了马尔可夫链状态转移图中转移概率的作用,并基于此思想提出了三种不同的隐写方法。首先提出了基于状态转移二进制序列的文本信息隐写方法。在该方法中,我们不仅紧紧围绕马尔可夫链中的转移概率这一概念对隐写模型进行设计,而且我们在信息通讯过程中对部分重要信息进行二次加密,进一步加强了信息的通讯安全。其次,我们提出了在模型中选择最优及次优语言序列的基于单比特规则的文本信息隐写方法,旨在将文本中较优的语言特征保存下来指导隐写文本的生成。最后,我们提出了一种基于变比特规则的文本信息隐写方法,在最大程度上保留了训练文本的语言特征,旨在实现最佳的文本隐写效果。我们对所提出方法完成了建模并选用了不同类别的数据进行实验,验证了方法的有效性。(3)开展了基于改进评价指标的文本信息隐写方法研究。我们在进一步的研究过程中发现,仅仅依据马尔可夫链状态转移图中的转移概率对文本的质量进行评价不够客观和全面。在一般文本中,语言序列的长度越长,该语言序列的出现频度将呈现下降的趋势。这些长语言序列或许拥有很好的语言特征及质量,仅仅由于其出现频度较低就将其从语言模型中删除,这对于模型质量的影响是负面的。因此,我们提出了基于改进评价指标和单比特规则的文本信息隐写方法。该方法通过建立新的语句质量评价标准和模型建立基准,旨在更加客观的对文本的质量加以评判,以期实现更好的性能。我们对所提出方法完成了建模并选用了不同类别的数据进行实验,验证了方法的有效性。论文在最后对于所做的研究进行了总结,并对以后的工作方向及内容做以展望。

【Abstract】 With the rapid development of information science and industry dominated by computer science and network technology,information security problems such as eavesdropping,replay,denial,information leakage,and illegal use are becoming increasingly prominent.Compared with the traditional information encryption technology,steganography can protect important information in a more covert way.Steganography is usually done with a carrier with a large amount of redundancy,such as images,audio,video,etc.As the oldest and most used media,the text is difficult to carry too much secret information due to its small redundancy,but it is also more valuable for research.In this thesis,we take natural language text as the research object of steganography and study the steganography methods of natural language text from different perspectives.The specific research contents and achievements are as follows:(1)The text information steganography method based on variable grammar rules is studied.In our research,we found that different dimensional grammar models capture different features of the training text during the construction process,which affects the quality of the final steganography text.Therefore,we combine different dimensional grammar models to take advantage of the advantages of different dimensional grammar models.In the research,we proposed two different steganography models of mixed grammatical rules based on this idea,namely,the text information steganography method based on half frequency crossover rule and the text information steganography method based on multi-rule language models alternation.Both of these methods try to complete the generation of steganography text based on more flexible language relations,hoping to achieve better steganography effect.We have completed the modeling of the proposed method and selected different types of data for experiments to verify the effectiveness of the method.(2)The text information steganography method based on transition probability isstudied.In the research,we highlight the role of transition probability in Markov chain state transition diagram and propose three different steganography methods based on this idea.Firstly,the text information steganography method based on state transition binary sequence is proposed.In this method,we not only design the steganography model around the concept of transition probability in Markov chain but also encrypt some important information in the process of information communication to further strengthen the communication security of information.Secondly,we propose the text information steganography method based on the single bit rule to select the optimal and suboptimal language sequences in the model,aiming to preserve the better language features in the text to guide the generation of steganography text.Finally,we propose the text information steganography method based on the variable bit rule,which preserves the linguistic features of the training text to the greatest extent,to achieve the best text steganography effect.We have completed the modeling of the proposed method and selected different types of data for experiments to verify the effectiveness of the method.(3)The text information steganography method based on improved evaluation index is studied.In the process of further research,we find that it is not objective and comprehensive to evaluate the quality of text only based on the transition probability in Markov chain state transition diagram.In general texts,the longer the length of the language sequence,the lower the frequency of the language sequence.These long language sequences may have good language features and quality,and they are deleted from the language model just because of their low frequency of occurrence,which has a negative impact on the quality of the model.Therefore,we propose the text information steganography method based on improved evaluation index and single bit rule.This method can evaluate the quality of text more objectively and achieve better performance by establishing a new evaluation standard of sentence quality and the establishment benchmark of model.We have completed the modeling of the proposed method and selected different types of data for experiments to verify the effectiveness of the method.At the end of the thesis,the research work is summarized,and the future work direction and content are prospected.

  • 【网络出版投稿人】 兰州大学
  • 【网络出版年期】2021年 04期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络