节点文献
数据库模式匹配系统研究
The Research on Database Schema Matching System
【作者】 钱颖;
【导师】 刘国华;
【作者基本信息】 燕山大学 , 计算机应用技术, 2006, 硕士
【摘要】 随着信息化时代的不断发展,对发掘异构模式之间语义一致性的要求日益迫切。模式匹配作为模式操作的第一步,在数据集成、数据转换、模型管理等领域都起到了非常关键的作用。本文对国内外关于模式匹配的研究现状进行了综合分析,从一个全新的角度对复杂的模式匹配方法进行了研究。首先,介绍了现有的两种复杂模式匹配方法。通过全面分析这两种方法中各处理单元的功能及应用的相关技术,指出了两种方法各自的优点和存在的不足。其次,基于以上的研究成果和现有的解决方案,提出了扩展的复杂模式匹配系统CSM。它可通过预处理和聚类处理分别从数据类型和数据值上过滤掉部分不合理的侯选匹配,并在匹配产生器中利用多个具有专门目的的检索程序对候选匹配空间的特殊部分进行检索,发掘1:1和复杂匹配;应用相似度估价器对候选匹配进行估价并由匹配选择器选择出最优的候选匹配;针对被匹配模式中存在不透明列的情况,还可进一步应用补充匹配器发掘不透明列间的匹配关系。从而发掘出全面、合理的匹配对。再次,根据全集进行模式匹配的思想,提出新的模式匹配方法——基于全集的复杂模式匹配。针对被匹配模式中缺乏匹配信息的问题,应用已知模式和模式间映射的全集为被匹配模式增添信息,提高匹配的查全率和查准率。最后,应用实例分析了上述两种模式匹配方法的匹配过程,并分别从理论和实验上证明了两种方法的可行性及有效性。
【Abstract】 With the development of information age, the demand of discovering semantic consistency between different schemas has been imminence increasingly. Schema matching as the first step of schema operation, it has a crucial effect in many fields, such as Data Integration, Data Translation and Model Management. This paper analyzes the current situation of the domestic and international schema matching problem, and researches for the problem of complex schema matching from a completely new perspective.At first, the two kinds of existing complex schema matching methods are introduced. This paper points out the merits and demerits in each method through analyzing the function of each process module and the interrelated technologies in the two methods in an all-round way.Secondly, on the basis of the above research and existent solutions, the extended complex schema matching system CSM is proposed. It filters some unreasonable matches on data types and values by preprocess and clustering process, and employs a set of special-purpose search procedures in match generator to explore a specialized portion of the search space and discovers 1:1 and complex matches; It estimates candidate matches and selects optimal candidate matches by using similarity estimator and match selector respectively; According to the problem that there are opaque columns in the schemas being matched, it can apply complementary matcher to find matching relations between opaque columns further more. Thereby it can discover more general, reasonable matching pairs.Moreover, depending on the idea of using a corpus to accomplish schema matching, a new schema matching method named corpus-based complex schema matching is proposed. According to the problem of the lack ofsufficient evidence in the schemas being matched, it can use a known corpus of schemas and mappings between schemas to augment the evidence about the schemas being matched, thereby improves the matching recall and precision.Finally, we analyze the matching process of the two schema matching methods as above using instances, and prove their feasibility and validity by theory and experiment.
【Key words】 Schema Matching; Complex Matching; Clustering; Similarity; Mutual Information; Corpus;
- 【网络出版投稿人】 燕山大学 【网络出版年期】2007年 02期
- 【分类号】TP311.13
- 【被引频次】5
- 【下载频次】182