节点文献

数据清理中同体不同源数据的数化算法研究

Digitization Algorithm Study of SEDS in Data Cleaning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 夏骄雄徐俊吴耿锋

【Author】 XIA Jiaoxiong,XU Jun,WU Gengfeng(School of Computer Engineering and Science,Shanghai University,Shanghai 200072)

【机构】 上海大学计算机工程与科学学院上海大学计算机工程与科学学院 上海200072上海200072

【摘要】 在数据仓库构建的数据清理过程中,同体不同源数据的发现一直是清理过程的难点。在现实情况下,存在的单一实体在不同的数据源中以不同的方式进行存储或者表达的同体不同源数据,传统数据清理技术对其发现、修正需要花费大量的时间和系统资源进行比较,实际效果并不理想。该文提出一种新型的、利用数据数字化存储特点来查找同体不同源数据的算法,能够有效减少数据间的比较次数,并确保数据清理结果的质量。

【Abstract】 It is always the difficulty to find out the “same entity from different sources(SEDS)” data in the data cleaning process of the data warehouse.The SEDS data are the same real world entities represented or stored differently in different data sources.The traditional data cleaning method costs a lot of system resources on finding and correcting such data,while the result is not ideal.With the digitization storage of the data,a new algorithm is proposed to find out the SEDS.The algorithm can reduce the comparison among the data effectively,and keep the quality at the same time.

【基金】 上海市高等学校科学技术青年基金资助项目(01QN59);上海市高等学校科学技术发展基金资助项目(04AB29)
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2007年01期
  • 【分类号】TP311.13
  • 【被引频次】9
  • 【下载频次】206
节点文献中: 

本文链接的文献网络图示:

本文的引文网络