节点文献

基于n-gram中英文字符串分割算法实现

Implementation of Algorithm Based on n-gram Chinese-English String Segmentation

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 何晓明洪亲蔡坚勇林鸿

【Author】 HE Xiao-ming,HONG Qin,CAI Jian-yong,LIN Hong(College of Photonic and Electronic Engineering of Fujian Normal University Cangshan Campus,Fuzhou 350007,China)

【机构】 福建师范大学仓山校区光电与信息工程学院

【摘要】 相似字符串的模糊查询是信息检索的重要组成部分,一直是人们研究的热点。目前基于关键词的查询技术都是前缀匹配,无法查找到与搜索字符串相似的结果。该文提出一种基于n-gram的中英文字符串分割技术的算法,该技术主要是对字符串进行中英文识别,然后基于n-gram按照指定长度进行分割,该技术是实现基于关键词的模糊查询技术的基础。该技术在数据清洗以及学位论文TMLC系统和垃圾邮件过滤等方面也有重要的应用前景。

【Abstract】 Similar string of fuzzy query is an important part of the information retrieval,has been the hotspot of the research.The keyword search technology is the prefix matching,unable to find similar results with the search string.This paper presents a n-gram based in the Chi nese-English string segmentation algorithm,the technique is mainly to string recognition based on n-gram in Chinese-English,then in ac cordance with the specified length of segmentation,the technique is realized based on keywords fuzzy query technology based.The tech nology in data cleaning and dissertations TMLC system and spam filtering has important application prospect.

【基金】 福建省自然科学基金项目(2010J01324)
  • 【文献出处】 电脑知识与技术 ,Computer Knowledge and Technology , 编辑部邮箱 ,2012年23期
  • 【分类号】TP391.3
  • 【被引频次】3
  • 【下载频次】162
节点文献中: