节点文献
基于SNM算法的大数据量中文商品清洗方法
Large Amount of Data in Chinese Commodity Cleaning MethodBased on the SNM Algorithm
【摘要】 SNM算法即邻近排序算法,是英文数据清洗最常用的算法[1]。目前为止,因为中英文语义的差异等一些原因,中文数据清洗还未形成完整的理论,现有中文数据清洗算法大多数是基于改编英文数据清洗算法而来的[2~3]。论文介绍数算法,论述该算法的缺陷,针对缺陷进项改进,并提出实际中的应用场景。通过实验结果显示,在相似重复记录消除方面,SNM改进算法具有明显的优势。
【Abstract】 SNM algorithm namely adjacent sorting algorithms,data cleaning is the most commonly used algorithm in English date cleaning. But so far,because of some reasons of the difference of semantics in both English and Chinese,Chinese datacleaning has not formed the perfect theory,most of the existing Chinese data cleaning algorithm is based on English data cleaning algorithm. This article will gradually introduce data cleaning,and will focus on based on the application of Chinese data cleaning SNMalgorithm. In this paper,the traditional algorithm of SNM is first introduced,the shortcomings of the algorithm are discussed. Thedefects are improved and the practical application scenarios are proposed. Comparing traditional SNM method and improved SNM algorithm through the experiment,the results show that in terms of similar duplicate records to eliminate,SNM improved algorithmhas obvious advantages.
- 【文献出处】 计算机与数字工程 ,Computer & Digital Engineering , 编辑部邮箱 ,2019年03期
- 【分类号】TP311.13
- 【被引频次】6
- 【下载频次】294