节点文献
基于词义特性的电子邮件敏感信息过滤仿真
Simulation of Email Sensitive Information Filtering Based on Word Meaning Characteristics
【摘要】 针对电子邮件敏感信息特征种类多,敏感信息过滤难度大的问题,提出一种基于决策树的过滤算法优化方法。建立电子邮件向量空间模型,给出信息对应词和所属类别向量关系,计算敏感信息中某一代表性词语与类别间的对应关系,通过词频出现概率求得所属类别,提取邮件特征。考虑到敏感信息在不同时间点的词义特性会发生变化,建立决策树,通过映射得到敏感信息与上下文信息串之间的影响关系,对电子邮件中的敏感信息项添加标签,求得属性值参数,按照参数大小设定邮件抗体的成熟度值,用于调整邮件传输通道宽度,实现精准过滤。实验数据证明,所提方法过滤精准度高,所需运算代价小,具有一定的实用价值。
【Abstract】 In order to address the problem of multiple types of sensitive information features in email and the high difficulty of filtering sensitive information, a method of optimizing the filtering algorithm was proposed based on decision tree. Firstly, we built a vector space model of email, and provided the vector relationship between the equivalent and categories of information. And then, we calculated the relationship between a representative word and a category in sensitive information, thus obtaining the category through the occurrence probability of word frequency, and extracting email features. After considering that the semantic characteristics of sensitive information might change at different points in time, a decision tree was constructed to map the impact relationship between sensitive information and context information strings. Moreover, sensitive information items in e-mails were labeled. Meanwhile, attribute parameters were obtained. The maturity value of the e-mail antibody was set according to the parameter size, which was used to adjust the width of the transmission channel. Finally, accurate filtering was achieved. Experimental data show that the proposed method has high filtering accuracy, low computational cost, and certain practical value.
【Key words】 Decision tree; E-mail; Sensitive information filtering; Maturity value; Context information string;
- 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2023年10期
- 【分类号】TP391.1;TP393.098
- 【下载频次】10