节点文献
多敏感属性数据发布隐私保护机制
Privacy Preservation Mechanism for Multiple Sensitive Attributes in Data Publishing
【作者】 王雪;
【导师】 姚琳;
【作者基本信息】 大连理工大学 , 软件工程(专业学位), 2023, 硕士
【摘要】 生活中数据的共享与发布往往涉及多个属性,而多个属性特别是敏感属性之间又往往存在着各种各样的关联关系,这些关系加剧了关系型数据的复杂性,如果直接发布,很可能会发生隐私泄露。为了保护发布数据的隐私,目前已经提出了许多隐私模型,然而这些模型大多优先考虑数据的隐私而不是可用性。因此,在发布多敏感属性数据的同时如何妥善处理属性间的关联以及实现可用性和隐私性的均衡等问题成为了当前的研究热点。本文深入对生活中包含多维敏感属性的关系型数据进行了研究,根据数据来源将其分为单源多敏感属性数据和多源多敏感属性数据,并分别针对这两种情况的数据发布提出了相应的隐私保护方案:针对单源多敏感属性数据,本文提出一种可用性感知的(α,β)隐私模型,同时,为了指导数据发布者合理设置α和β,本方案设置了隐私增益和效用损失的度量,并定量地用隐私来置换效用,反之亦然。此外,本文使用提升度指标量化敏感属性(SAs)与敏感属性之间的关联,使用卡方值来量化准标识符(QIs)与敏感属性之间的关联,并基于它们应用抑制和排列两种匿名化技术来适当地匿名化它们。针对多源多敏感属性数据,本文改进了单源方案中对SA-SA和QI-SA两种相关关系的划分和处理,提出一种基于联邦学习的多源隐私方案,分别使用横向联邦和纵向联邦来处理水平数据和垂直数据的多源联合发布,依据对关联关系的训练结果,分别应用抑制和排列技术来适当地匿名化它们并保证隐私和可用性的均衡,与此同时,本文提出一种新颖的联邦终止机制,通过代入对相关性程度的考量,进一步提升模型性能。本文选取真实数据集并使用Python语言环境对上述两种隐私保护方案进行了仿真实验,并与现有的代表性研究在数据的辨别性、记录隐匿率、附加信息损失、隐私增益以及可用性损失等方面进行了对比分析。理论分析和实验结果显示,上述两种隐私方案与现有工作相比能够更为全面地抵御现有种类的隐私泄漏,且均能够在实现高数据隐私的同时保证数据更好的可用性。
【Abstract】 The sharing and publishing of data in life often involves multiple attributes,and there are often various association relationships between multiple attributes,especially sensitive attributes.These relationships aggravate the complexity of relational data.If published directly,privacy disclosure is likely to occur.In order to protect the privacy of data publishing,many privacy models have been proposed.However,most of these models give priority to data privacy rather than utility.Therefore,how to properly handle the association between attributes and achieve the tradeoff of utility and privacy while publishing multiple sensitive attribute data has become a current research hotspot.This paper deeply studies the relational data containing multiple dimensional sensitive attributes in life,and divides it into single-source multiple sensitive attribute data and multisource multiple sensitive attribute data according to the data source.And then this paper proposes the corresponding privacy protection schemes for data publishing in these two cases:For single-source multiple sensitive attribute data,this paper proposes a utility-aware(α,β)privacy model.At the same time,in order to guide the data publisher to reasonably set α andβ,this scheme sets the measurement of privacy gain and utility loss,and quantitatively trades privacy for utility and vice versa.In addition,this paper uses the lift degrees to quantify the association between sensitive attributes(SAs)and sensitive attributes and the chi-square values to quantify the association between quasi-identifiers(QIs)and sensitive attributes with applying two anonymization techniques of suppression and permutation to appropriately anonymize them.For multi-source multiple sensitive attribute data,this paper improves the division and processing of SA-SA and QI-SA association in the single-source scheme,and proposes a multisource privacy scheme based on federated learning,which uses horizontal federation and vertical federation to handle the joint publishing of horizontal data and vertical data respectively.According to the training results of the association relationship,the suppression and permutation techniques are applied respectively to appropriately anonymize them and ensure the balance of privacy and utility.At the same time,this paper proposes a novel federation termination mechanism,which takes into account the degree of association to further improve the performance of the model.This paper selects real data sets and uses Python language environment to carry out simulation experiments on the above two privacy preservation schemes,and makes a comparative analysis with existing representative studies in the aspects of data indiscernibility,record suppression ratio,additional information loss,privacy gain and utility loss.The theoretical and experimental results show that the above two privacy schemes can more comprehensively resist the existing types of privacy disclosures than the existing work and achieve high data privacy while ensuring better data utility.
【Key words】 Privacy Preservation; Multiple Sensitive Attributes; Privacy-utility Tradeoff; Association Concealment; Federated Learning;
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2025年 07期
- 【分类号】TP309;TP181