节点文献

基于马尔可夫聚类的隐私高维数据发布方法

Private high-dimensional data publication with Markov clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘卓群龙士工张珺铭刘光源

【Author】 LIU Zhuo-qun;LONG Shi-gong;ZHANG Jun-ming;LIU Guang-yuan;State Key Laboratory of Public Big Data, Guizhou University;College of Computer Science and Technology, Guizhou University;Foundation Department,Guizhou Polytechnic of Construction;

【通讯作者】 龙士工;

【机构】 贵州大学公共大数据国家重点实验室贵州大学计算机科学与技术学院贵州建设职业技术学院基础部

【摘要】 针对现有差分隐私的方法在处理高维数据发布时面临计算成本高、数据精度低和中心服务器不可信任的问题,提出一种基于马尔可夫聚类的隐私高维数据发布方法MCL-LDP。基于在用户本地实现对用户数据的隐私保护,中心服务器接收到用户本地化差分隐私保护的数据后,构建无向依赖图矩阵表示高维数据的复杂的属性关联性,基于马尔可夫聚类将高维数据属性集分割成多个低维属性簇,利用EM算法计算低维属性簇和重叠属性簇的边缘分布、估计原始数据的联合分布,通过采样合成新的数据集进行发布。实验结果表明,所提出方法在发布高维数据集上有较好的精度、较少的迭代次数和较高的计算效率。

【Abstract】 Aiming at the problems of high computing cost, low data accuracy and untrustworthy central server in the existing differential privacy methods, a privacy high-dimensional data publishing method based on Markov clustering, namely MCL-LDP, was proposed. Based on the realization of privacy protection of user data locally, after receiving the data of localized differential privacy protection of users, the central server constructed an unoriented dependency graph matrix to represent the complex attribute correlation of high-dimensional data. The attribute set of high-dimensional data was split into multiple low-dimensional attribute clusters based on Markov clustering. The EM algorithm was used to calculate the edge distribution of low-dimensional attribute clusters and overlapping attribute clusters to estimate the joint distribution of the original data. A new data set was synthesized by sampling and published. Experimental results show that the proposed method has better accuracy, less iterations and higher computational efficiency on publishing high-dimensional data sets.

【基金】 国家自然科学基金项目(62062020)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2025年01期
  • 【分类号】TP309
  • 【下载频次】20
节点文献中: 

本文链接的文献网络图示:

本文的引文网络