节点文献

基于类间离散度的文档敏感内容识别算法研究

The Algorithm of Recognizing Sensitive Documents’ Content Based on Discrete Degree betweenClass

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 秦艺文杨榆

【Author】 Qin Yiwen;Yang Yu;Information Security Center of Beijing University of Posts and Telecommunications;

【机构】 北京邮电大学信息安全中心

【摘要】 敏感数据信息一旦被外泄,后果将不堪设想。而防泄密管理中亟待解决的重大问题,即是如何能快速、准确地从大量数据信息识别敏感内容。本文首先基于敏感文本库,训练已知分类文本集;在简便有效的文本敏感特征提取方法的基础上,引入类间离散因子修正传统的TF-IDF权值确定方法;随后利用支持向量机构建分类器,以识别和判断敏感文本内容。实验表明,在查准率、查全率、F1测试值,虚警、漏检,以及处理时间等方面,该算法具有较高的准确性和高效性。

【Abstract】 Once the sensitive information was leaked,the result will be unimaginable.The first major problem to be solved in management of leak prevention is how to quickly and accurately identify sensitive content from a large amount of data.In this paper,firstly training the set of known classification texts based on the sensitive text library.Following a simple and effective text sensitive feature extraction method,the paper introduces the discrete factor between class in order to correcting TF- IDF weight determining equation.Finally building a classifier using support vector machine to identify sensitive contents.The results show that from the point of precision rate,recall rate,F1test value,false alarm,the missing rate and the processing time,this algorithm has higher accuracy and efficiency.

  • 【会议录名称】 第十届中国通信学会学术年会论文集
  • 【会议名称】第十届中国通信学会学术年会
  • 【会议时间】2014-09-05
  • 【会议地点】中国辽宁沈阳
  • 【分类号】TP309
  • 【主办单位】中国通信学会、辽宁省通信管理局
节点文献中: