节点文献

大数据发布隐私保护技术研究

Research on Privacy Preserving Big Data Publishing Technology

【作者】 晏燕

【导师】 郝晓弘;

【作者基本信息】 兰州理工大学 , 控制理论与控制工程, 2018, 博士

【摘要】 移动互联网的快速发展和智能终端的广泛使用,使得个人信息数字化程度不断提高,促进了大数据时代的到来。大数据在推动各行业技术发展、提高数据资源服务能力等方面发挥了巨大作用,但同时也给个人隐私安全带来了严峻挑战。大数据的多源异构和动态发布特性增强了不同数据源之间的关联性,容易导致隐私信息的泄露和隐私保护方法的失效,不但损害用户的名誉、财产和生命安全,甚至威胁到国家信息安全。因此,对大数据发布隐私保护技术的研究是关系到大数据安全应用和进一步开发使用的重要环节。本文针对大数据发布环节的隐私保护问题,分析了大数据发布过程中的隐私致险因素和不确定性;设计了大数据发布隐私风险态势评估指标体系和评价方法,建立起大数据发布“预警系统”;提出简化的大数据关联表示方法和准标识符属性判定算法,为隐私保护操作确定了关键属性集合;针对静态和动态大数据的不同特点分别设计了相应的大数据隐私保护发布算法,相比现有同类算法在隐私保护效果和算法性能方面具有显著提高。全文研究工作包括以下几个方面:(1)大数据发布隐私风险态势评估技术研究。结合大数据发布的特点及应用模式,定义了大数据发布环境下的隐私风险、隐私资产、隐私威胁和隐私脆弱性,构建了三级隐私风险态势评估指标体系。基于集对分析理论设计了隐私风险态势评估的方法,提出基于多元偏联系数的最小二偏赋权方法,消除了隐私风险态势评估过程中不确定性因素对指标权重分配的干扰和影响。实例分析和对比实验表明,基于集对多元偏联系数的隐私风险态势评估方法较好地体现了风险指标的状态和发展变化趋势,能够实现对大数据发布系统隐私风险状态和风险因素的动态跟踪及评估。(2)大数据关联表示与准标识符属性判定方法研究。针对多源异构大数据关联性隐蔽且复杂、容易导致发布数据隐私泄露的问题,设计了基于图的大数据实体关联表示方法,并进一步抽象为链接待发布数据与已发布数据和外部知识的属性图。分析并定义了属性图中准标识符的作用,从集合的独立性角度出发将准标识符属性的判定问题转化为求解属性图的割点问题,并进一步设计了基于割点的准标识符属性判定算法,为阻止链接攻击实现隐私保护操作确定了关键属性集合。与现有其他准标识符属性求解方法的对比分析表明,本文提出的基于割点的准标识符属性判定算法具有更好的划分效果和更低的计算复杂度。(3)静态大数据隐私保护模糊发布技术研究。针对传统k-匿名隐私保护模型计算复杂度高、信息损失度大、k值难以准确设定等问题,提出基于模糊语义的静态大数据转换发布方法。根据数值型和分类型敏感属性的不同特点,分别设计了基于集对云模型和语义泛化树的模糊发布算法;建立了模糊语义区分度和泛化信息保留度参量,较好的反映出发布信息与原始信息之间的区别与联系。通过在阿里云平台上实际大数据集的运行比较,表明本文算法比其他模糊语义和聚类保护方法具有更低的计算复杂度和更好的发布数据可用性。(4)动态位置大数据差分隐私划分发布技术研究。针对位置大数据动态统计发布存在的索引结构和隐私预算不确定性问题,提出基于差分隐私保护模型的分层混合划分发布算法。采用均匀时间间隔内连续发布数据快照的方法对动态位置大数据进行采样平滑处理;通过密度自适应网格划分实现不同采样时刻位置大数据的空间聚类;设计了基于区域均匀性的启发式四叉划分算法和相应的隐私分配策略,不但解决了自顶向下空间划分时难以确定停止条件的问题,同时均衡了噪声误差和均匀假设误差对发布数据查询精度的影响。通过在阿里云平台上实际位置大数据集的运行比较,表明本文算法在改善范围查询精度和算法运行效率方面具有较大的优势。

【Abstract】 The rapid development of Mobile Internet and the widespread use of smart terminals lead to continuous increase in the digitization of personal information,which promotes the arrival of the era of big data.Big data plays an significant role in promoting technology development in various industries and improving service capabilities of data resources.However,it also brings serious challenges to personal privacy security.The multi-source heterogeneous and dynamic publishing features of big data enhanced correlationship between different data sources,which easily lead to the disclosure of private information and the failure of privacy protection methods.These will not only damage users’ reputation,property and life safety,but even threaten the national information security.Therefore,the research on privacy preserving data publishing is an important part related to security application and further developments and use of big data.This dissertation aims at the issue of privacy protection during the process of big data publishing,and analyzes the privacy risk factors and uncertainties.System of indicators and evaluation methods of privacy risk situation assessment are designed to form "pre-warning system" for the privacy preserving publishing of big data.A simplified big data association representation method as well as the quasi-identifier attribute identification algorithm are proposed,which help to determine the set of key attributes for privacy protection operation.According to different characteristics of static big data and dynamic big data,corresponding privacy preserving publishing algorithms are designed.Compared with some existing similar algorithms,the privacy protection effect and algorithm performances are significantly improved.The research work of the full dissertation includes the following aspects:(1)Research on privacy risk situation assessment of big data publishing.Combined with the characteristics and application modes of big data release,privacy risks,privacy assets,privacy threats and privacy vulnerabilities are defined for the environment of big data publishing.A three-level privacy risk situation assessment index system is established.Then,a privacy risk situation assessment method is designed based on the theory of set pair analysis,and a least squares partial weighting method is proposed based on the partial connection numbers,which can eliminate the interference and influence of uncertainty factors on allocation of weighting.Case analysis and comparison experiments show that the proposed privacy risk situation assessment method better reflects the status and development trend of privacy risk indicators,and can track and evaluate the privacy risk status and risk factors of the big data release system dynamically.(2)Research on relevance representation of big data and identification of quasi-identifier attributes.Aiming at the problem that multi-source heterogeneous big data have concealedand complex relevance and easily lead to the disclosure of privacy,a graph-based big data entity association representation method is designed,which is further abstracted into attribute graph by linking the publishing data with published data and external knowledge.The roles of quasi-identifiers within attribute graph have been analyzed and defined.The problem of determining quasi-identifier attributes is converted into the problem of finding cut-vertexes for attribute graph from the perspective of the independence of set.Further more,a quasi-identifier partitioning algorithm is designed based on cut-vertex,which determined the set of key attributes to prevent linking attacks and implement privacy protection operations.Compared with some existing quasi-identifier identification methods,the proposed algorithm has better partitioning effect and lower computational complexity.(3)Research on fuzzy publishing technology for privacy protection of static big data.Aiming at the problems of traditional k-anonymity model such as computational complexity,large information loss and difficult problem of determining k value,a fuzzy semantics publishing method for static big data is proposed.The fuzzy publishing algorithm based on set-pair cloud model and the semantic generalization tree are designed according to different characteristics of numerical sensitive attributes and categorical sensitive attributes.Parameters such as fuzzy semantic distinction and generalized reserve degree are designed to reflect the relationship between the published information and the original information.Comparison examinations carried out on Ali cloud platform show that the proposed algorithm has lower computational complexity and better availability of published data than other privacy preserving fuzzy semantics and clustering algorithms.(4)Research on differential privacy decomposition technology for privacy protection of dynamic location big data.Aiming at the uncertainty problems of spatial index structure and privacy budget allocation in dynamic publishing of statistical location big data,a hierarchical hybrid decomposition algorithm is proposed based on differential privacy model.The method of continuously publishing data snapshots on average time interval is used to sample and smooth the dynamical location big data.Spatial clustering of location big data snapshots on different sampling times are carried out by the adaptive density grid partitioning algorithm.Heuristic quad-tree partitioning method based on regional uniformity as well as corresponding dynamic privacy allocation strategy were proposed,which not only solved the problem of determining stop condition for top-down space decomposition,but also equalized the impact of noise error and uniform hypothesis error on query accuracy after publishing.Comparison examinations carried out on Ali cloud platform show that the proposed algorithm has larger advantages in improving regional query accuracy and the operating efficiency.

节点文献中: