节点文献
面向半结构化文本的领域本体关系抽取
Automatic Domain-Ontology Relation Extraction from Semi-Structured Texts
【Author】 Cheng Xiao~1,Zheng Dequan~1,Yang Yuhang~1,Shao Guojun~2 1 Department of Computer Science and Technology,Harbin Institute of Technology,Harbin 150001 2 Beijing Shougang Automatic Information Technology Co.,Ltd,Qinhuangdao 066004
【机构】 哈尔滨工业大学计算机科学与技术学院; 秦皇岛首信自动化系统工程有限公司;
【摘要】 本文提出了一种以半结构化文本作为数据源,进行领域本体关系抽取的方法。首先,利用概念实例和属性值的共现得到文档集合。其次,定义关系模式形式,从文档集合中得到关系模式实例,包括关系模式实例的聚类以及类内合并。最后,将各类关系模式用于抽取领域本体新实例的属性值信息。在针对电影,图书和音乐三个领域进行的实验中,关系模式聚类的错分率和漏分率分别为0.19%,1.31%,类内合并后关系模式的准确率最高可达85%。实验结果表明了本方法对于领域本体中关系抽取的有效性。
【Abstract】 This paper presents a new method to acquire Domain-Ontology relations from semi-structured data sources. First,obtain documents according to the co-occurrence of concept instance and attribute value.Further,define formats of relation patterns,and extract pattern instances from documents,including pattern clustering and pattern combining in each cluster.Finally,relation pattern instances are applied to gain attribute values of new concept instances in Domain-Ontology.Experiments are carried out in fields of film,book and music,the rate of pattern incorrect -division and pattern leakage are respectively 0.19%and 131%,the highest precision of combined relation patterns reaches 85%. Experimental results demonstrate that the method developed in this paper is fairly efficient.
【Key words】 semi-structure; Domain-Ontology; relation extraction; pattern instance;
- 【会议录名称】 中国计算机语言学研究前沿进展(2007-2009)
- 【会议名称】第十届全国计算语言学学术会议
- 【会议时间】2009-07-24
- 【会议地点】中国山东烟台
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会