节点文献

一种基于聚类树的增量式数据清洗算法

An incremental algorithms of data cleansing based on clustering tree

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘芳何飞

【Author】 Liu Fang He Fei Liu Fang Dr.; College of Computer Sci. & Tech., Huazhong Univ. of Sci.& Tech., Wuhan 430074, China.

【机构】 华中科技大学计算机科学与技术学院华中科技大学计算机科学与技术学院 湖北武汉430074湖北武汉430074

【摘要】 研究了在数据模式与匹配规则不变的前提下 ,数据集动态增加时近似重复记录的识别问题 ,提出了一种基于聚类树的增量式数据清洗算法IACT .该算法通过构建聚类树先对记录进行分区 ,然后在划分的区域内进行相似度的计算识别出近似重复记录 ,从而完成了增量式相似重复记录的检测 .实验结果证明了IACT算法在无损精度的情况下 ,在效率上优于多趟邻近排序 (MPN)算法 .

【Abstract】 This paper studied the problem of detecting approximately duplicate records while receiving increments of data with no changes in data schema and matching rule set, and presented an incremental algorithm IACT (Incremental Algorithms based on Clustering Trees for data cleansing). IACT divided the data records into a few areas and computed their similarity to identify the approximately duplicate records to accomplish the data cleansing task in the partitioned areas through creating clustering tree. Compared with the algorithm MPN, the experimental result proves that IACT algorithm is more effective while possessed of the same precision.

【基金】 国家“十五”重大科技基金资助项目 (2 0 0 1BA10 2A0 6 11) .
  • 【文献出处】 华中科技大学学报(自然科学版) ,Journal of Huazhong University of Science and Technology , 编辑部邮箱 ,2005年03期
  • 【分类号】TP311
  • 【被引频次】11
  • 【下载频次】409
节点文献中: 

本文链接的文献网络图示:

本文的引文网络