节点文献
基于语义的Deep Web数据源自动发现技术
An Automatic Discovery Technology of Deep Web Data Sources Based on Semantic
【Author】 FANG Wei, HU Peng-yu, ZHAO Peng-peng, CUI Zhi-ming (Institute of Intelligent Information Processing and Application, SooChow University, Suzhou 215006, China)
【机构】 苏州大学智能信息处理及应用研究所;
【摘要】 为了方便用户快捷高效的使用DeepWeb中内容丰富、主题专一的高质量信息,对DeepWeb数据源发现研究已成为一个非常迫切的问题。目前通用的方法是基于关键词的主题过滤策略,这样容易发现一些不相关的数据源,为此提出一种新的基于语义的DeepWeb数据源聚焦爬行方法,利用朴素贝叶斯分类算法自动发现DeepWeb数据源,实验验证了该方法的有效性。
【Abstract】 To expediently utilize the rich ,oriented topic and high quality information of Deep Web, this problem on Deep Web data sources discovery has been focused by more and more people. Nowadays, topic filtering strategy based on
【关键词】 Deep Web;
语义;
本体;
表单;
【Key words】 is widely used, then it will obtain some irrelevant data sources. This paper proposes a new focused crawling method based on semantic for Deep Web data sources, and describes a technique for detecting query interface using naive Bayes classification. Finally, the method is validated by test. Key words: Deep Web; Semantic; Ontology; Form; Bayes Classification;
【Key words】 is widely used, then it will obtain some irrelevant data sources. This paper proposes a new focused crawling method based on semantic for Deep Web data sources, and describes a technique for detecting query interface using naive Bayes classification. Finally, the method is validated by test. Key words: Deep Web; Semantic; Ontology; Form; Bayes Classification;
【基金】 国家自然科学基金项目(60673092);2005年度教育部科研重点项目(205059);教育部高校博士学科点科研基金(20040285016);江苏省高技术研究计划项目(BG2005019)
- 【会议录名称】 2007年全国开放式分布与并行计算机学术会议论文集(上册)
- 【会议名称】2007年全国开放式分布与并行计算机学术会议
- 【会议时间】2007-10-12
- 【会议地点】中国广西南宁
- 【分类号】TP391.3
- 【主办单位】中国计算机学会开放系统专业委员会