节点文献

云环境下异构信息交换模板的研究与设计

Heterogeneous in for Mationexc Hange Templates Research and Design of Cloud Environment

【作者】 贾伟

【导师】 杨文川;

【作者基本信息】 北京邮电大学 , 通信与信息系统, 2012, 硕士

【摘要】 在日益海量的信息数据环境下,云计算理论与技术成为处理大规模信息数据的主要手段。云计算的核心思想,是将大量用网络连接的计算资源统一管理和调度,构成一个计算资源池向用户按需服务。国家某部委在这样的云环境背景下,提出了异构信息交换模板的探究设计的课题,本文基于在与其合作课题下完成了如下研究工作:1.在对异构信息交换模板理论研究的基础上,研究与设计了异构信息交换模板系统的数据输入与解析。所谓异构信息交换模板这个概念需要分拆解释,异构顾名思义,就是指整体结构互不相同,信息交换是目标,当异构的文本被标准化成统一的模板时,就能达到数据共享,也就是互相间的信息补充、交换。此模块主要是将多种异构的数据源信息纳入系统并解析成树结构文档对象模型。由于课题输入数据源类别多,数量大,文档对象模型也不同。主要的输入数据为基于HTML的网络页面信息、Adobe pdf文档以及微软的MS-WORD文件,不同数据源采用了不同的解析策略。2.在解析好树状文档对象模型的基础上,研究与设计了如何对数据进行筛滤、提取后,构建起信息交换模板雏形。数据过滤前后,以准确率、召回率、F值等指标进行性能跟踪。在模板构建时,采用工厂模式与单例模式等,面向接口的设计模块功能,保证了更高的可扩展性。3.在模板构建后,根据模板属性,对其进行信息交换模板归类并序列化存储于数据库,根据反馈不断调整参数优化系统。对于模板的归类原则,应用贝叶斯分类器原理来进行归类。在序列化存储时使用JAVA序列化API,结合数据库连接池技术存于数据库中。根据反射机制,不断提高模板构建的准确率、召回率及F值等性能。在课题研究与设计中,在云环境背景下,对异构信息交换模板系统进行了分模块分析,系统设计以及系统实现。根据时间占用度,空间占用度及准确率、查全率、F值等进行了性能评测,系统良好。

【Abstract】 In the increasingly vast amounts of information and data environment, cloud computing theory and large-scale information technology is as the main data processing means. The core idea of cloud computing is to a large number of network connections with unified management and scheduling of computing resources, constitute a pool of computing resources on demand service to users. Ministry of Industry and environment in the context of this cloud, made a template of heterogeneous information exchange designed to explore the subject, this subject is based on their cooperation to complete the following studies:First, the theory of heterogeneous information exchange research is based on the template, research and information exchange designed template heterogeneous data entry and analysis system. The so-called heterogeneous information exchange need to be separately explained the concept of templates, heterogeneous by definition, refers to the overall structure different from each other, exchange of information is the goal, when the heterogeneous text is standardized into a single template, can achieve data sharing, but also is the complement of information between each otherand exchange. This module is mainly to multiple heterogeneous data sources into the system and parse the information into a document object model tree. As the subject of the input data source categories and large quantities, document object model is different. The main input data for the HTML-based web page information, Adobe pdf documents and Microsoft’s MS-WORD documents, different data sources using different analytical strategies.Second, a good tree in the parse document object model is based on the research and design of how the data sieve, extracted, from information exchange to build a template shape. Data before and after filtration, the accuracy, recall, F value and other indicators of performance tracking.When building the template, using the factory pattern with a single case, the mode module for interface design features, to ensure a higher scalability.Third, build the template, based on the template properties, its exchange of information classification and sequence of the template stored in the database, based on feedback continuously adjusted to optimize system parameters. Principles for the classification of the template, the application of principles of Bayesian classifier to classify.Stored in the serialized serialized using JAVA API, combined with database connection pool stored in the database. According to reflection, to continuously improve the accuracy of the template to build the recall rate and F-value properties.In research and design, in the context of the cloud environment, the exchange of information on the heterogeneous system was sub-module template analysis, system design and system implementation. According to the time occupancy, occupancy degree and accuracy, recall, F value for the performance evaluation, etc., the system is good.

【关键词】 云环境异构信息模板java
【Key words】 CloudHeterogeneousInformationTemplateJava
节点文献中: 

本文链接的文献网络图示:

本文的引文网络