节点文献

基于并列结构的概念实例和属性的同步提取方法

To Extract Concept Instances and Concept Attributes Based on Coordinate Structure

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李文杰穗志方

【Author】 LI Wenjie1,2,SUI Zhifang1,2(1.Institute of Computational Linguistics,Peking University,Beijing 100871,China; 2.Key Laboratory of Computational Linguistics(Ministry of Education),Peking University,Beijing 100871,China)

【机构】 北京大学计算语言学研究所北京大学计算语言学教育部重点实验室

【摘要】 在概念实例和属性的提取研究中,针对基于模式的方法召回率比较低的特点,该文提出了一种基于并列结构的概念实例和属性的同步提取方法。首先利用并列结构模式去网页集合中提取同类词语集合,然后再用基于种子的弱指导方法去学习实例和属性共现的上下文模式,最后再通过模式去提取候选实例或候选属性。在此过程中,每提取出一个候选,就将该候选所在的同类词语集合合并到候选集合中。实验结果表明,该文的方法在不降低准确率的基础上,能大大提高提取结果的召回率。

【Abstract】 Most researches on concept instances and concept attributes extraction focuses on pattern-based approaches,which usually suffer from a low recall rate.In this paper,we present a method of extracting concept instances and concept attributes based on the coordinate structure.Since a part of candidate instances and attributes extracted by the coordination patterns can be putted into the similar-concept-phrases sets in advance,we can use these similar-concept-phrases sets to expand the extraction results in the procedure of co-occurrence pattern-based extraction.Compared with the baseline without using the coordination patterns,experimental results show that the coverage of this method is significantly improved without reducing the precision.

【基金】 国家自然科学基金(60873156、61075067);国家社会科学基金(09BYY032)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2012年02期
  • 【分类号】TP391.7
  • 【被引频次】9
  • 【下载频次】146
节点文献中: 

本文链接的文献网络图示:

本文的引文网络