节点文献

基于统计的中文地名识别

Identification of Chinese Place Names Based on Statistics

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄德根岳广玲杨元生

【Author】 HUANG De gen YUE Guang ling YANG Yuan sheng(Department of Computer Science and Engineering,Dalian University of Technology,DaLian 116024,China)

【机构】 大连理工大学计算机科学与工程系大连理工大学计算机科学与工程系 大连116024大连116024大连116024

【摘要】 本文针对有特征词的中文地名识别进行了研究。该系统使用从大规模地名词典和真实文本语料库得到的统计信息以及针对地名特点总结出来的规则 ,通过计算地名的构词可信度和接续可信度从而识别中文地名。该模型对自动分词的切分作了有效的调整 ,系统闭式召回率和精确率分别为 90 2 4 %和 93 14 % ,开式召回率和精确率分别达 86 86 %和 91 4 8%。

【Abstract】 Unknown word recognition is one of the challenging tasks in natural language processing research.This paper proposes a place name identification model in dictionary based Chinese word segmentation,in which we used statistical information drawn from a training corpus to calculate lexical reliability and contextual reliability.The rules of Chinese place names are also used in the model.We applied this approach to a Chinese morphological analysis system,and achieved 90.24% recall and 93 14% precision in close test,while the recall and precision also reach 86 86% and 91 48% in open test.

【基金】 国家自然科学基金资助项目 (6 0 14 30 0 2 )
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2003年02期
  • 【分类号】TP391.4
  • 【被引频次】167
  • 【下载频次】1140
节点文献中: