节点文献

基于全集的复杂模式匹配

Corpus-Based Complex Schema Matching

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 钱颖聂俊岚刘国华郜时红

【Author】 Qian Ying,Nie Junlan,Liu Guohua,Gao Shihong (College of Information Science and Engineering,Yanshan University,Qinhuangdao 066004)

【机构】 燕山大学信息科学与工程学院

【摘要】 模式匹配是数据集成和数据转换中的重要问题.现有的模式匹配方法大多集中于发掘模式间的1:1匹配,然而,在现实世界模式之间除了1:1匹配还包括很多的复杂匹配.提出一种基于全集的复杂模式匹配方法,它可应用模式和映射的全集为被匹配模式添加信息;然后,利用多个具有特殊目的的检索程序分别对候选空间的特殊部分进行检索,发掘1:1和复杂匹配;最后通过学习全集中元素及元素间关系的统计,自动推导出可过滤候选匹配的约束,生成最优的匹配.实验表明,该方法不仅能全面地发掘模式间匹配,与其他复杂模式匹配方法相比,还具有较高的查全率和查准率.

【Abstract】 Schema matching is a crucial problem in data integration and data translation.The existing approaches to automating schema matching focus on computing direct element matches(1:1 matches) between two schemas.However,relationships between real-world schemas involve many complex matches besides 1:1 matches.A corpus-based complex schema matching method is introduced in this paper.It uses a corpus of schemas and mappings to augment the evidence about the schemas being matched.Then,it employs a set of special-purpose searchers to explore a specialized portion of the search space and discovers 1 :1 and complex matches.Finally,it learns statistics about elements and relationships in corpus and uses them to infer constraints that can be used to prune candidate mappings and generates optimal matches. Experiments show that this method does not discover matches between schemas roundly,but also improves the recall and precision compared with other complex schema matching methods in practice.

【关键词】 模式匹配全集增广模型统计
【Key words】 schema matchingcorpusaugmented modelstatistic
【基金】 教育部科学技术研究重点基金项目(205014))
  • 【会议录名称】 第二十三届中国数据库学术会议论文集(研究报告篇)
  • 【会议名称】第二十三届中国数据库学术会议
  • 【会议时间】2006-11-10
  • 【会议地点】中国广东广州
  • 【分类号】TP311.13
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: