节点文献

基于规则抽取的中文时间表达式识别技术研究

Research on Chinese Time Expression Recognition Technology Based on Rule Extraction

【作者】 张昊

【导师】 胡伟;

【作者基本信息】 南京大学 , 计算机技术(专业学位), 2018, 硕士

【摘要】 时间表达式识别是自然语言处理领域中,命名实体识别技术的一个重要组成部分。时间表达式的识别和对时间信息的获取和使用,在信息检索、自动问答等诸多领域有着重要的作用。本文围绕中文时间表达式识别这一具体问题,首先介绍时间相关基本概念和时间识别基本方法,而后具体阐述时间识别各方法的特点,并提出新的时间表达式识别方法,主要工作和贡献如下:1、提出了一种基于规则抽取的时间表达式识别方法,并提出了规则评价指标。利用训练数据集中时间表达式用词和词性的特点,自动地从训练集中抽取识别规则,避免大量人工构建识别规则的过程,使用规则评价指标,对抽取出的识别规则进行评价和筛选,然后使用规则集识别时间表达式。2、提出了一种基于规则抽取结合机器学习分类算法进行筛选的时间表达式识别新方法。基于规则抽取的时间表达式识别方法可以在抽取了足够多规则的情况下获得较高的召回率,但是由于规则本身没有充分利用上下文信息和语义信息,导致识别正确率不高。结合机器学习分类算法,通过训练二元分类器,对规则识别出的候选时间表达式进行过滤,从而在保证召回率的情况下,大幅提高识别正确率。在微软亚洲研究院中文命名实体识别语料上使用本文提出的基于规则抽取和分类算法的时间表达式识别新方法,取得了较好的识别效果,F1值可达94.05%,优于基于CRF和基于规则的方法。

【Abstract】 Time expression recognition is an important part of named entity recognition in natural language processing.Recognizing time expressions and acquiring and using time information play a very important role in many fields,such as information retrieval,question answering,etc.This thesis focuses on a concrete problem,Chinese time expression recognition.First,time related basic concepts and basic methods to recognizing time expressions are introduced.Then,features of these methods are expounded in detail and a new method to recognizing time expressions is proposed.The main work and contributions of this thesis are as follows1.A method based on rule extraction to recognize time expressions and criteria to evaluate rules are proposed.Automatically extracting rules that can recognize time expressions by using word and POS features of time expressions in training data can avoid many manual work in process of constructing rules.Then,criteria is used to evaluate and filter extracted rules and the processed rule set is used to recognize time expressions.2.A new method that recognize time expressions based on rule extraction and use classification algorithms in machine learning to filter expressions is proposed.Methods based on rule extraction can achieve relatively high recall rate with enough rules,but they cannot achieve high precision rate because context and semantic information are not used sufficiently by rules themselves.With classification algorithms added in,training classifiers using training data and using them to filter candidate expressions can raise precision rate while recall rate will keep relative high.Using new method based on rule extraction on Microsoft Research Asia corpus for Chinese named entity recognition to recognize time expressions can achieve relatively good result that the F1 score can reach 94.05%,which is better than the method based on CRF and the method based on rules.

  • 【网络出版投稿人】 南京大学
  • 【网络出版年期】2021年 01期
  • 【分类号】TP391.1;TP181
  • 【被引频次】2
  • 【下载频次】70
  • 攻读期成果
节点文献中: