节点文献
基于n-gram中英文字符串分割算法实现
Implementation of Algorithm Based on n-gram Chinese-English String Segmentation
【摘要】 相似字符串的模糊查询是信息检索的重要组成部分,一直是人们研究的热点。目前基于关键词的查询技术都是前缀匹配,无法查找到与搜索字符串相似的结果。该文提出一种基于n-gram的中英文字符串分割技术的算法,该技术主要是对字符串进行中英文识别,然后基于n-gram按照指定长度进行分割,该技术是实现基于关键词的模糊查询技术的基础。该技术在数据清洗以及学位论文TMLC系统和垃圾邮件过滤等方面也有重要的应用前景。
【Abstract】 Similar string of fuzzy query is an important part of the information retrieval,has been the hotspot of the research.The keyword search technology is the prefix matching,unable to find similar results with the search string.This paper presents a n-gram based in the Chi nese-English string segmentation algorithm,the technique is mainly to string recognition based on n-gram in Chinese-English,then in ac cordance with the specified length of segmentation,the technique is realized based on keywords fuzzy query technology based.The tech nology in data cleaning and dissertations TMLC system and spam filtering has important application prospect.
【Key words】 fuzzy query; n-gram; string segmentation; edit distance; data mining;
- 【文献出处】 电脑知识与技术 ,Computer Knowledge and Technology , 编辑部邮箱 ,2012年23期
- 【分类号】TP391.3
- 【被引频次】3
- 【下载频次】162