节点文献

基于条件随机场的中文科研论文信息抽取

Information Extraction from Chinese Research Papers Based on Conditional Random Fields

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 于江德樊孝忠尹继豪

【Author】 Yu Jiang-de Fan Xiao-zhong Yin Ji-hao(School of Computer Science and Tech.,Beijing Institute of Tech.,Beijing 100081,China)

【机构】 北京理工大学计算机科学技术学院北京理工大学计算机科学技术学院 北京100081北京100081

【摘要】 科研论文头部信息和引文信息对基于域的论文检索、统计和引用分析是必不可少的.由于隐马尔可夫模型不能充分利用对抽取有用的上下文特征,因此文中提出了一种基于条件随机场的中文科研论文头部和引文信息抽取方法,该方法的关键在于模型参数估计和特征选择.实验中采用L-BFGS算法学习模型参数,并选择局部、版面、词典和状态转移4类特征作为模型特征集.在信息抽取时先利用分隔符、特定标识符等格式信息对文本进行分块,在分块基础上用条件随机场进行指定域的抽取.实验表明,该方法抽取性能明显优于基于隐马尔可夫模型的方法,且加入不同的特征集对抽取性能提升作用不同.

【Abstract】 The information of headers and citations of research papers is necessary for many applications,such as the field-based paper search,the paper statistics and the citation analysis.In order to enhance the utilization of context features for information extraction which is greatly restricted by the hidden Markov model(HMM),a method based on the conditional random fields(CRFs) is proposed to extract the information of paper header and citation from Chinese research papers.The proposed method,whose key is the parameter estimation and the feature selection,employs L-BFGS algorithm for the estimation of model parameters in the experiment and selects the categories features of location,layout,lexicon and state transition as the feature set of the model.During the information extraction,the format information about list separators and special-labels is used to segment the text,and then CRFs are applied to the extraction in special fields.Experimental results show that the proposed method possesses better performance than that based on the HMM,and that the performance improvement varies with the features sets.

【基金】 教育部博士点基金资助项目(20050007023)
  • 【文献出处】 华南理工大学学报(自然科学版) ,Journal of South China University of Technology(Natural Science Edition) , 编辑部邮箱 ,2007年09期
  • 【分类号】TP391.1
  • 【被引频次】41
  • 【下载频次】768
节点文献中: 

本文链接的文献网络图示:

本文的引文网络