节点文献

面向生物文本的实体关系自动抽取问题研究

Research on Automated Biomedical Relation Extraction from Bio-literature

【作者】 张宏涛

【导师】 朱小燕;

【作者基本信息】 清华大学 , 计算机科学与技术, 2012, 博士

【摘要】 生物实体之间的关系是当前生物学家尤为关注的知识之一。然而,大量的实体关系蕴藏在海量的生物文本中,并且随着生物文本的快速增长而持续累积。如何从生物文本中自动地抽取实体关系已经成为生物信息学领域的一个重要挑战。本文以面向生物文本的实体关系自动抽取为主要研究对象,围绕实体关系自动抽取过程中面临的特征构造、数据集不平衡、标注数据集规模较小、跨领域标注数据集复用等关键问题展开研究。论文工作包括:(1)针对特征构造问题,提出了一种紧凑的特征向量。该特征向量具有两大优点:一是融合了多种特征,包括词、词性、句法以及词模板等在内的多类特征,信息丰富并能对复杂的生物文本进行有效表达;二是具有紧凑的特征表示方式,较好的缓解了因融合丰富特征导致的特征稀疏问题。(2)针对数据集不平衡问题,提出了基于自适应的欠采样方法和基于动态联合学习的随机欠采样方法。基于自适应的欠采样方法在欠采样过程中能对分类器进行自适应的调整,而基于动态联合学习的随机欠采样方法融合了过采样和欠采样的思想,实现了在扩大的样本空间上进行欠采样。它们均有效降低了利用欠采样思想解决数据集不平衡问题所致的删除有益样本的风险。(3)针对标注数据集规模较小的问题,提出了统一的主动学习框架。它是一种更适用于实体关系抽取任务的主动学习框架。除样本选择模块外,它进一步融入了多样性样本选择模块、主动特征获取模块以及相关特征选择模块。实验结果表明,该框架有效降低了抽取方法对标注数据集规模的依赖程度。(4)针对跨领域标注数据集复用问题,建立了基于迁移学习的复用框架,具体包括基于样本迁移学习的复用方法、基于特征组迁移学习的复用方法以及基于主动学习和迁移学习融合的复用方法。其中,基于样本迁移学习的复用方法和基于特征组迁移学习的复用方法分别从样本和特征两个粒度进行迁移学习,实现了跨领域标注数据集的复用;而基于主动学习和迁移学习融合的复用方法则融合了主动学习和迁移学习的优点,为解决更实际的问题奠定了基础。

【Abstract】 The important relations between biomedical entities have attracted significantattention from biologists. However, an enormous number of biomedicalrelations are buried in millions of biomedical research articles that have beenpublished over the years, and the number is growing. Rediscovering them frombio-literature automatically is a challenging bioinformatics task. In this thesis,we focus on some key issues in the biomedical relation extraction task,including the construction of feature vector, the class imbalance problem indatasets, the small-labeled datasets and the re-use of labeled datasets fromdifferent domains. The major contributions of this thesis are as follows:(1) For the construction of feature vector, we propose a compact feature vector.The two major advantages of this feature vector are its rich features and its compactfeature representations. Specifically, it integrates keyword features, part-of-speechfeatures, syntactic features and lexical pattern features, which can express muchimportant information for the complex biomedical text; it employs the compact featurerepresentations for the above features, which can alleviate the sparse feature problemcaused by the integration of rich features.(2) For the class imbalance problem in datasets, we propose two methods based onunder-sampling, i.e., the adaptation based under-sampling method and the dynamicco-training based random under-sampling method. The former method makes anadaptive adjustment for the classifier during the under-sampling process, while the lattermethod enables under-sampling on an expanded sample space by integrating theunder-sampling and over-sampling. Both methods reduce the risk of losing criticalsamples when employing under-sampling based methods to address the class imbalanceproblem.(3) For the small-labeled datasets, we propose a unified active learningframework. The proposed framework is a more appropriate active learningframework for the biomedical relation extraction task. In addition to thecommon data selection module, it integrates the diverse data selection module,the active feature acquisition module and the informative feature selection module. The experimental results show that the proposed framework effectivelyreduces the reliance on the size of labeled datasets.(4) For the re-use of labeled datasets from different domains, we build are-use framework based on transfer learning. Specifically, the re-use frameworkincludes the instance based transfer learning re-use method, the feature groupbased transfer learning re-used method and the integration of active learning andtransfer learning re-used method. The former two methods accomplish there-use of cross-domain labeled datasets by transfer learning based on instancegranularity and feature granularity, while the latter method integrates theadvantages of active learning and transfer learning, in order to shed some lighton more practical tasks.

  • 【网络出版投稿人】 清华大学
  • 【网络出版年期】2013年 07期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络