节点文献

异构复杂信息网络敏感数据流动态挖掘

Dynamic mining of sensitive data streams in heterogeneous complex information networks

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 熊菊霞吴尽昭

【Author】 XIONG Ju-xia;WU Jin-zhao;Chengdu Institute of Computer Application,Chinese Academy of Sciences;University of Chinese Academy of Sciences;Guangxi Key Laboratory of Hybrid Computational and IC Design Analysis,Guangxi University for Nationalities;

【机构】 中国科学院成都计算机应用研究所中国科学院大学广西民族大学广西混杂计算与集成电路设计分析重点实验室

【摘要】 针对异构复杂信息网络中存在高维冗余的敏感数据流,可挖掘数据特征形成概率较低,导致需要多次挖掘、挖掘内存占用高、挖掘精度低、时间长的问题,提出基于最大类间散度的网络敏感数据流动态挖掘方法。将敏感数据的差异最大化间隔作为分类基础,得到网络敏感数据的最大类间散度,在遗传迭代状态下确定最优散度迭代函数,对迭代函数进行挖掘特征优选,得出动态可挖掘特征。对可挖掘特征进行聚类分析,挖掘得到数据隐藏信息模式,并对其进行评价,将合理的信息模式进行知识表示,从而实现异构复杂信息网络敏感数据流动态挖掘。实验结果表明,所提方法可挖掘特征形成概率高达98%,labels标记与实际值较为接近。所提方法挖掘精度高,且运行时间较短、内存占用率低。

【Abstract】 For the sensitive data streams with high-dimensional redundancy in heterogeneous complex information networks, the probability of data feature formation is low, which leads to multiple mining, high memory usage, low mining accuracy and long running time. Aiming at the above problems, a dynamic network sensitive data stream mining method based on the maximum inter-class divergence is proposed. The maximum difference interval between sensitive data is used as the basis for classification to obtain the maximum inter-class divergence of the network sensitive data. The optimal divergence iterative function is determined in the genetic iterative state. The mining characteristics of the iterative function are preferably selected to obtain the dynamic mining characteristics. Clustering analysis is performed on the mining characteristics to obtain data hiding information modes. These modes are evaluated, and knowledge representation is carried out on the reasonable information modes, so as to realize the dynamic mining of the sensitive data streams in the heterogeneous complex information networks. The experimental results show that the mineable feature formation probability of the method can be up to 98%, and the labels are close to the actual values. The method has the advantages of high mining accuracy, short running time and low memory usage.

【基金】 国家自然科学基金(61772006);广西科技重大专项项目(AA17204096);广西科技基地和人才专项项目(2016AD05050);广西“八桂学者”专项资助;广西高校中青年教师基础能力提升项目(2017KY0174)
  • 【文献出处】 计算机工程与科学 ,Computer Engineering & Science , 编辑部邮箱 ,2020年04期
  • 【分类号】TP311.13;TP181
  • 【被引频次】8
  • 【下载频次】152
节点文献中: 

本文链接的文献网络图示:

本文的引文网络