节点文献

基于SNM算法的大数据量中文商品清洗方法

Large Amount of Data in Chinese Commodity Cleaning MethodBased on the SNM Algorithm

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张苗苗苏勇

【Author】 ZHANG Miaomiao;SU Yong;School of Computer,Jiangsu University of Science and Technology;

【机构】 江苏科技大学计算机学院

【摘要】 SNM算法即邻近排序算法,是英文数据清洗最常用的算法[1]。目前为止,因为中英文语义的差异等一些原因,中文数据清洗还未形成完整的理论,现有中文数据清洗算法大多数是基于改编英文数据清洗算法而来的[2~3]。论文介绍数算法,论述该算法的缺陷,针对缺陷进项改进,并提出实际中的应用场景。通过实验结果显示,在相似重复记录消除方面,SNM改进算法具有明显的优势。

【Abstract】 SNM algorithm namely adjacent sorting algorithms,data cleaning is the most commonly used algorithm in English date cleaning. But so far,because of some reasons of the difference of semantics in both English and Chinese,Chinese datacleaning has not formed the perfect theory,most of the existing Chinese data cleaning algorithm is based on English data cleaning algorithm. This article will gradually introduce data cleaning,and will focus on based on the application of Chinese data cleaning SNMalgorithm. In this paper,the traditional algorithm of SNM is first introduced,the shortcomings of the algorithm are discussed. Thedefects are improved and the practical application scenarios are proposed. Comparing traditional SNM method and improved SNM algorithm through the experiment,the results show that in terms of similar duplicate records to eliminate,SNM improved algorithmhas obvious advantages.

【关键词】 SNM算法数据清洗重复记录
【Key words】 SNM algorithmdata cleaningduplicate records
  • 【文献出处】 计算机与数字工程 ,Computer & Digital Engineering , 编辑部邮箱 ,2019年03期
  • 【分类号】TP311.13
  • 【被引频次】6
  • 【下载频次】294
节点文献中: 

本文链接的文献网络图示:

本文的引文网络