节点文献

中文命名实体识别及其关系抽取研究

The Research of Chinese Named Entity Recognition and Its Relation Extraction

【作者】 温锐

【导师】 朱巧明;

【作者基本信息】 苏州大学 , 计算机应用技术, 2005, 硕士

【摘要】 本文主要研究了中文命名实体识别及其关系抽取,设计和实现了一个能识别和抽取人名、地名和机构名的系统CNEE,并通过SRV算法实现了个人主页中的人名和E-mail 的抽取。CNEE 先进行自动分词和词性标注,然后根据人名、地名和机构名的各自特点来进行识别和抽取。自动分词和词性标注的准确率将会直接影响命名实体的识别。本文改进了词性标注,在HMM 标注的基础上引入负反馈规则来进行修正,改进后的词性标注准确率在96%。实验表明CNEE 抽取人名、地名和机构名的F 指数均达到了75%以上。SRV 算法是一个基于规则学习的关系抽取算法,具有训练样本少和准确率高的优点。本文将SRV 算法用于个人主页中的人名和Email 的抽取,取得了较好的效果。实验证明SRV 算法用于命名实体关系的抽取是成功和有效的。

【Abstract】 The paper mainly researches the Chinese named entity recognition and its relation extraction, designs and realizes a Chinese named entity extraction (CNEE) system, which can recognize and extract person, location and organization name, and extracts person name and E-mail from personal homepage by the SRV algorithm. CNEE first automatically segments words and tags parts of speech, then recognizes and extracts person, location and organization name according to their different traits. The paper improves the POS tagging by introducing the negative feedback rules into HMM tagging. The final precision of POS tagging is about 96 percents, while the final experiment shows that the F-measures of the person, location and organization name extraction of CNEE are all over 75 percents. SRV is an algorithm based on rule learning, which needs only a few train samples and has high precision. The paper applies it into extracting person name and Email from personal homepage and obtains good result. The experiment demonstrates that SRV is successful and effective for the named entity relation extraction.

  • 【网络出版投稿人】 苏州大学
  • 【网络出版年期】2006年 04期
  • 【分类号】TP391.4
  • 【被引频次】13
  • 【下载频次】936
节点文献中: 

本文链接的文献网络图示:

本文的引文网络