节点文献

中文信息检索引擎中的分词与检索技术

Word Segment and Search Techniques for Chinese Information Search Engines

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吴栋滕育平

【Author】 WU Dong, TENG Yu ping(Laboratory of Pure Mathematics and Combinatorics, Center for Combinatorics, Nankai University, Tianjin 300071, China)

【机构】 南开大学组合数学研究中心核心数学与组合数学教育部重点实验室南开大学组合数学研究中心核心数学与组合数学教育部重点实验室 天津300071天津300071

【摘要】 文中论述了在开发中文信息检索系统中所涉及到的两项关键技术 ,即中文分词技术和检索技术。针对中文分词技术 ,介绍了一种改进的正向最大匹配切分算法 ,以及为消除歧义引入的校正策略 ,并在此基础上结合统计方法处理未登录词。针对检索技术 ,综述了几种最常用的检索模型的原理 ,并对每种模型的优缺点进行了简要分析。最后对给出的分词算法进行了测试 ,测试结果表明该分词算法准确度和效率能够满足实用的要求

【Abstract】 Two key techniques in the development of Chinese Information Retrieval System are discussed in this paper, i.e., Chinese word segmentation and search technique. For Chinese word segmentation, the paper presents an improved MM segmentation algorithm, the revise strategy for disambiguation, and the statistic method for unknown words recognition based on the previous methods. For search technique, the paper summarizes the principle of several kinds of search models, and analyzes the advantages and disadvantages of each model simply. At last, the given segmentation algorithm is evaluated, and the results reveal that the veracity and efficiency of the algorithm can satisfy the applied request.

  • 【文献出处】 计算机应用 ,Computer Applications , 编辑部邮箱 ,2004年07期
  • 【分类号】TP391.3
  • 【被引频次】193
  • 【下载频次】1646
节点文献中: