节点文献

条件随机场与领域本体元素集相结合的未登录词识别研究

The Study on Out-of-Vocabulary Identification on a Model Based on the Combination of CRFs and Domain Ontology Elements Set

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 段宇锋朱雯晶陈巧刘伟刘凤红

【Author】 Duan Yufeng;Zhu Wenjing;Chen Qiao;Liu Wei;Liu Fenghong;Business School, East China Normal University;Shanghai Library;School of Public Economics and Administration, Shanghai University of Finance and Economics;Institute of Botany, Chinese Academy of Sciences;

【机构】 华东师范大学商学院上海图书馆上海财经大学公共经济与管理学院中国科学院植物研究所

【摘要】 【目的】建立未登录词识别模型,提升发现自然科学领域文本中未登录词的能力,同时降低人工干预成本。【方法】在假设的基础上,构建条件随机场(CRFs)与领域本体元素集相结合的未登录词识别模型。以生物多样性文本为样本,通过比较不同模型性能的差异,检验假设,验证模型的合理性。【结果】实验结果表明,CRFs模型选择单纯的字、字词混合序列、字词混合序列及默认词性、字词混合序列及含自定义语义功能标记的词性为特征时,未登录词识别能力依次提升。该结果证明研究假设为真,本文建立的模型科学、合理。【局限】模型标注未登录词的准确性有待提升。【结论】该模型具有更强的未登录词识别能力,同时可以极大地降低人工建立训练集的成本。

【Abstract】 [Objective] Establish a model to improve the out-of-vocabulary identification capability, reduce the cost of manual intervention. [Methods] On the basis of the hypothesis, a out-of-vocabulary identification model is set up combining CRFs and domain Ontology elements set. Using biodiversity text as samples, the rationality of the model is verified by comparing the performance differences among models and testing hypothesis. [Results] The experimental results show that the model established by this study has the best identification capability. The results prove that the hypothesis is true, and the model is reasonable and scientific. [Limitations] The tagging accuracy of the model remains to be improved. [Conclusions] The model established in this paper has better identification capability, while greatly reducing the cost of artificial training dataset.

【基金】 国家社会科学基金一般项目“基于无监督语义标注的网络中文学术信息抽取研究”(项目编号:11BTQ024)的研究成果之一
  • 【文献出处】 现代图书情报技术 ,New Technology of Library and Information Service , 编辑部邮箱 ,2015年04期
  • 【分类号】TP391.1
  • 【被引频次】10
  • 【下载频次】275
节点文献中: 

本文链接的文献网络图示:

本文的引文网络