节点文献
基于中文科技论文的本体交互式构建方法研究
Research on Chinese Scientific and Technical Paper Based Interactive Ontology Building
【作者】 张新;
【导师】 党延忠;
【作者基本信息】 大连理工大学 , 系统工程, 2006, 硕士
【摘要】 同行评议是科学基金项目评审工作的核心,其效果主要取决于同行专家的选择。本质上,同行专家的选择过程是在已知项目知识的条件下,从专家库中搜索出与己知项目具有相似知识的专家的过程,也可以看作是一个语义检索的过程。本体是语义检索的基础和核心,本体质量的好坏直接影响到语义检索的效果。同时,本体的手工构建由于耗时费力也严重阻碍了本体大规模的应用。因此,本体的自动构建是一个亟待解决的问题。 本文在对国内外本体自动构建相关研究进行全面分析和总结的基础上,提出了基于中文科技论文的本体交互式构建方法。该方法基于系统集成创新的思想,充分利用现有的自然语言处理技术和统计学习方法,从特定领域内的自然语言文本中提取领域概念以及概念间的语义关系。本文的核心工作包括以下三点: (1) 领域概念的提取。主要是通过基于长度递减与串频统计的文本切分方法以及汉语短语词法规则,提取领域的候选概念,然后通过统计方法分析领域归属度并基于词典进行概念约简,得到由多个词和短语组成的与领域相关的概念。 (2) 语义关系的获取。主要是通过关联规则挖掘、依存句法分析以及机器学习方法来学习表达语义关系的关系句法模式,应用已得到的句法模式析取语义关系,并对概念间的语义关系进行命名。 (3) 本体交互式构建原型系统的分析、设计和实现。系统主要分为三个模块:文本管理、本体构建和本体维护,其中重点介绍了本体构建模块的功能和具体实现。 在此基础上,本文以“计算机科学”及其子领域“计算机硬件”的文本为试验对象,基于本文提出的本体交互式构建方法和原型系统构建了一个小规模的领域本体,并对试验结果进行了分析。试验结果表明,本文提出的基于中文科技论文的领域本体构建方法具有较高的准确性,并且不依赖于领域词典,适用于任何领域本体的构建,具有较大的通用性,能够辅助领域专家更高效、准确地完成本体构建的任务。
【Abstract】 The peer review of grant proposals is important to academics from all disciplines, which depends on the selection of reviewers. Generally speaking, reviewers’ selection is to search experts who have similar knowledge with the proposals, in other words, it can be treated as a process of semantic retrieval, where ontology appears to be an important part. However, manually building ontologies turns to be time consuming and labor intensive, thus inhibits the applications. As a result, automatic ontology building is an urgent problem to be resolved.In this paper, theories about automatic ontology building are comprehensively analyzed and a novel method for Chinese scienfitic and technical papers based interactive ontology building is proposed. The method makes full use of ideas from system integrated innovation, natural language processing techniques and statistical algorithms to extract domain-specific concepts as well as their semantic relations from domain texts. Key work of this paper can be summarized in the following three parts:(1) Domain-specific concepts extraction. Here domain-specific concepts are extracted from texts through semantic string segmentation, rules matching, analysis of domain consensus and domain relevance, and concept reduction.(2) Semantic relations learning. First associated concepts are mined by association rule, and then sentences containing them are analyzed by dependency tree parser. Then syntactic patterns are generated by machine learning algorithms and are used to extract relations between new associated concepts. At last, relations are named according to the patterns.(3) The implementation of the interactive ontology building system. The system has three modules: text management, ontology building, and ontology maintenance, and ontology building module is strengthened and described in detail.In the end, a small ontology is built from Chinese texts in Computer Science domain and its sub domain Computer Hardware, and experimental results are evaluated, which show that the method for interactive ontology building proposed in this paper can improve the efficiency to a great extent and show the promise of relatively high precision. Furthermore, it is domain lexicon independent, thus is suitable for any domain ontology building.
【Key words】 Ontology; Ontology Building; Domain-specific Concept; Semantic Relation; Relation Syntactic Pattern;
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2007年 03期
- 【分类号】TP399
- 【被引频次】9
- 【下载频次】253