节点文献

一种基于N-Gram的检测相似重复记录的高效方法

An N-Gram Based Approach for Detecting Approximately Duplicate Database Records

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 邱越峰田增平季文赟周傲英

【Author】 Qiu Yuefeng,Tian Zengping,Ji Wenyun,Zhou Aoying (Department of Computer Science Fudan University Shanghai 200433)

【机构】 复旦大学计算机系

【摘要】 如何消除数据库中的重复信息已成为数据质量研究中一个热门话题.本文提出了一种基于N-Gram的检测相似重复记录的方法,主要工作有:(1)给出了一种高效的基于N-Gram的聚类算法,该算法能适应常见的拼写错误如插入、删除、替换、交换等,复杂度为O(N);(2)介绍了一种高效的应用无关的Pairwise比较算法,复杂度为O(k~2);(3)采用了一种改进的优先队列算法来准确地聚类相似重复记录.

【Abstract】 Eliminating duplications in large database has drawn some attentions.In this paper,the problem is studied.An N-gram based approach is introduced.The contributions of this paper are:(1) an efficient n-gram based clustering algorithm is given,which can tolerate the most common types of errors in misspelled words like insertion,deletion,transposition,substitution,and reordering of the words in a record.It is only with the computing complexity of O(N).(2) A very efficient application independent Pairwise comparison algorithm is presented.It has the computing complexity of O(k~2).(3) For detecting duplicate records in the indexed table,an algorithm that modifies the priority queue method is presented.

【关键词】 N-GramRNGNpairwise聚类优先队列
【Key words】 N-GramRNGNpairwiseclusteringpriority queue
  • 【会议录名称】 第十六届全国数据库学术会议论文集
  • 【会议名称】第十六届全国数据库学术会议
  • 【会议时间】1999-08-24
  • 【会议地点】中国甘肃兰州
  • 【分类号】TP311.13
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络