节点文献
基于本体相似度的中文科研论文信息抽取
Information Extraction from Chinese Research Papers Based on Ontology Similarity
【摘要】 随着大量的科研论文出现在互联网上,从中精确地抽取论文头部信息和引文信息显得十分重要。提出了基于本体相似度的信息抽取方法,该方法的关键在于用本体相似度判定某个行本体是正例还是反例,然后通过主动学习选择最有可能包含抽取信息的行本体集,再充分利用本体的语义推理能力找到正确的片断。从论文中提取头部信息和引文信息为进一步的语义检索和语义存储奠定基础。测试数据集的实验结果显示该方法比其他方法具有较高的准确率。
【Abstract】 Information extraction from Chinese research papers based on ontology similarity abstract as many research papers appear on the Internet,it becomes more and more important to extract paper header information and citations accurately from these papers.Presents a new information extraction algorithm which is based on ontology similarity.The key point of the algorithm is to divide the row-ontology samples into positive and negative instances,extract the most appropriate set of row-ontologes by active learning,and then retrieve the correct pieces lie in them by using the reasoning mechanism contained in the ontologies.It can get header information and citation from these papers,which assist the semantic searching and storage.Test results show that the algorithm is more precise than other approaches.
【Key words】 information extraction; ontology similarity; semantic reasoning; active learning;
- 【文献出处】 计算机技术与发展 ,Computer Technology and Development , 编辑部邮箱 ,2008年12期
- 【分类号】TP391.3
- 【被引频次】2
- 【下载频次】258