节点文献

基于贝叶斯方法的Web服务分类的研究

The Research of Web Services Classification Based on Bayesian Technology

【作者】 费爱蓉

【导师】 穆斌;

【作者基本信息】 合肥工业大学 , 计算机应用技术, 2004, 硕士

【摘要】 随着互联网的发展,当前出现的Web标准如WSDL,SOAP,UDDI,DAML-S,使得Internet成为一个异构的、具有互操作性的Web服务的海洋,从而使应用程序的开发过程简化为发现Web服务和集成Web服务的过程。 Web服务是基于XML的可被远程调用的网络组件,自动调用和集成Web服务的关键在于将机器可理解的语义元数据与Web服务关联起来。因为Web服务的描述缺乏足够的语义信息,Web服务发现具有不确定性,为了能够根据用户提供的信息更加准确地描述并执行Web服务,必须考虑更加丰富的语义和上下文信息。贝叶斯技术和贝叶斯网络是从传统的统计学中分离出来的,对不确定性问题进行处理的一个有力工具,它以完善的贝叶斯理论为基础,有较强的模型表示、学习和推理能力。本论文就是探索采用贝叶斯方法对Web服务自动生成语义元数据,来对Web服务进行分类以提高Web服务检索的效率。 本论文中基于贝叶斯技术的Web服务分类算法的主要思想是通过引入贝叶斯潜在语义模型,首先将含有潜在类别主题变量的WSDL文档分配到相应的类主题中;接着利用朴素贝叶斯模型,结合前一阶段的知识,完成对未含类主题变量的文档作标注。本算法不需要对大量训练样本的类别标注,只需提供相应的类主题变量,从而提高了WSDL文档分类的自动性和效率。实验结果表明,本算法具有较高的分类正确率。

【Abstract】 Emerging Web standards such as WSDL, SOAP, UDDI and DAML-S promise a network of heterogeneous yet interoperable Web Services. Web Services would greatly simplify the development of many kinds of data integration and knowledge management applications.Web Services are networked components that can be invoked remotely using standard XML-based protocols. The key to automatically invoking and composing Web Services is to associate machine-understandable semantic metadata with each service. A central challenge to the Web Services initiatives is therefore to construct tools to (semi-)automatically generate the necessary metadata. Bayesian technology and Bayesian networks have been successfully used to process artificial intelligence problems and to discover knowledge in databases domain. We explored the specific machine learning technique, i.e. Bayesian technique in this paper, in the hope of automatically creating such kind of metadata from training data.We assigned WSDL documents to different categories using Bayesian Latent Semantic model. With the frame of BLSM, our system classified WSDL documents only by a few of latent class variables and no labeled data. Two steps were included: the first step was to label those documents containing latent class variables by BLSA; the second step was to label the rest by Naive Bayesian model with EM algorithm. It has achieved good precision and proved to be very effective and efficient in our experimental system.

  • 【分类号】TP393.09
  • 【被引频次】14
  • 【下载频次】448
节点文献中: 

本文链接的文献网络图示:

本文的引文网络