节点文献

基于不完备集双聚类的缺失数据填补算法

Missing Data Filling Algorithm Based on Incomplete Set Biclustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 韩飞沈镇林

【Author】 HAN Fei;SHEN Zhenlin;College of Information Science and Technology,Jinan University;Office of Information Management,Jinan University;

【机构】 暨南大学信息科学与技术学院暨南大学信息管理办公室

【摘要】 缺失数据填补是数据清洗领域的一个重要问题。由于绝大部分局部填补方法基于全部属性进行分类,未考虑对象属性之间的关联性,因此基于不完备集双聚类,提出一种缺失数据填补算法。该算法利用双聚类完美簇的平均平方残基为0及簇内的属性值波动一致的特点,对缺失数据进行填补。通过数学分析,把寻找含有缺失值的最大完美簇问题转化为求解缺失对象与其他对象之间的最大相似属性集问题,在相同的最大相似属性集下,以缺失值的众数作为填补值。采用4组UCI数据集进行实验,结果表明,该算法相比ROUSTIDA算法平均提高了77.13%的填补值精确度。

【Abstract】 Missing data filling is an important issue in the field of data cleaning.As the vast majority of local filling methods realize classification on the basis of all attributes without considering the correlation between object attributes,this paper puts forward a missing data filling algorithm based on incomplete set biclustering.This algorithm fills missing data based on the theory that the mean squared residue of biclustering perfect cluster is O and the fluctuation of the cluster s attribute values is consistent.This paper translates the problem of finding the maximum perfect cluster which contains the missing values into the problem of finding out the maximum similarity attribute sets between the missing object and other objects through mathematical analysis,then the majority of missing values which is used as the filling value can be calculated by the same maximum similarity attribute sets.This paper takes experiments using 4 groups of UCI data sets,and it is demonstrated that the proposed algorithm averagely improves the accuracy of 77.13%filling values compared with ROUSTIDA algorithm.

【基金】 广东省高新技术产业化基金资助项目(2011B080701046)
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2016年04期
  • 【分类号】TP311.13
  • 【被引频次】24
  • 【下载频次】291
节点文献中: 

本文链接的文献网络图示:

本文的引文网络