节点文献

基于维基百科的中文嵌套命名实体识别语料库自动构建

Automatic Construction of Chinese Nested Named Entity Recognition Corpus Based on Wikipedia

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李雁群何云琪钱龙华周国栋

【Author】 LI Yanqun;HE Yunqi;QIAN Longhua;ZHOU Guodong;Natural Language Processing Laboratory,School of Computer Science and Technology,Soochow University;

【通讯作者】 钱龙华;

【机构】 苏州大学计算机科学与技术学院自然语言处理实验室

【摘要】 传统的监督学习方法需要标注一定规模的领域内语料库,限制了其领域适应性。为此,提出一种从中文维基百科条目中自动构建中文嵌套命名实体识别语料库的方法。对中文维基百科的条目进行实体分类,利用实体条目构造实体的嵌套结构,从而自动生成大规模的中文嵌套命名实体识别语料库。在手工标注嵌套命名实体识别语料库上的实验结果表明,自动构建的语料库具有规模较大、领域广的特点,且能够适应宽泛领域上的中文嵌套命名实体识别任务。

【Abstract】 Traditional supervised learning method needs to label the corpus in a certain scale,which limits its domain adaptability. Therefore,a method of automatically constructing a Chinese nested named entity recognition corpus from Chinese Wikipedia entries is proposed. The Chinese Wikipedia entries are classified into entities entries,and the nested structure of the entities is constructed by using the entity entries,thereby automatically generating a large-scale Chinese nested named entity recognition corpus. Experimental results on the manually labeled nested named entity recognition corpus show that the automatically constructed corpus has the characteristics of large scale and wide field, and can adapt to the Chinese nested named entity recognition task in a wide range of fields.

【基金】 国家自然科学基金(61373096,61331011,61673290)
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2018年11期
  • 【分类号】TP391.1
  • 【被引频次】12
  • 【下载频次】510
节点文献中: 

本文链接的文献网络图示:

本文的引文网络