节点文献

健康体检数据仓库的构建与分析系统的实现

Construction of Data Warehouse and Implementation of Analysis System for Health Examination

【作者】 孙喆

【导师】 卢朝霞;

【作者基本信息】 东北大学 , 计算机软件与理论, 2015, 硕士

【摘要】 随着健康体检业务的不断发展以及体检用户的不断增多,健康体检系统中积累了大量宝贵的数据。如何有效利用这些体检数据为医生和管理者提供决策支持成为相关机构面临的共同问题。本文针对此问题设计实现了一个健康体检数据分析系统。首先,引入数据仓库技术为健康体检数据分析提供了独立的环境,解决了健康体检数据存储和集成问题。通过健康体检数据仓库的维度建模过程,本文对健康体检数据建模涉及的事实、维度、粒度等进行了详细讨论。经过合理的模型构建,仓库数据被重新组织成了适于分析的结构。采用Shell和PL/SQL等高级脚本语言编码实现的ETL系统实现了每日数据的定时加载和更新,同时保证了最大的便捷性和灵活性。其次,为了实现健康体检数据的多维分析,使医生和管理者获得多角度分析关键指标的能力,本文引入了 OLAP技术。通过使用OLAP工具MSTR极大地简化了多维分析报表的开发。利用其提供的ROLAP服务器可以读取关系型数据仓库中的事实表和维度表,将相关数据表模型化成为一个统一的多维度模型。经过工具的配置可以定制多维模型中虚拟立方体的汇集计算结果,最终为医生和管理者提供健康体检数据多维分析报表服务。最后,本文探讨了健康风险评估的方法。通过引入数据挖掘技术中的分类技术,探索用户检验指标和检查结论之间的联系并建立相应的预测模型。文中选取决策树、朴素贝叶斯和支持向量机这三种常用的分类模型在真实数据上进行了实验,三种分类器的准确率都达到了 80%以上,证明了分类方法用于健康风险评估的可行性。此外,针对实验中健康体检数据集出现的非平衡性问题进行了讨论,最终选用数据预处理中的过采样方法对训练数据进行均衡。在对比实验中使用SMOTE算法对训练数据进行预处理之后,三种分类算法在总体分类准确率变化不明显的情况下对少数关注类的分类能力获得了显著提升,最终证明了过采样方法在健康体检数据集的不平衡性问题上应用的可行性。

【Abstract】 With the development of health examination and the increasing of users,a lot of valuable data are accumulated in the health examination system.It has become a common problem that the relevant agencies are facing to provide decision support for doctors and managers effectively using health examination data.To solve this problem,we design and implement a health examination data analysis system in this paper.Firstly,we provide an independent health examination analysis environment by using the data warehouse technology.Health examination data warehouse solves the data storage and integration issues.In this paper,we discuss the facts、the dimensions and the grains of the dimension model in health examination area in detail.After the dimensional modeling,the health examination data in warehouse is reorganized into a structure which is suitable for analysis.The ETL system which is programed with high-level scripting languages,such as Shell and PL/SQL,implements the daily data loading and updating,and guarantees the maximum convenience and flexibility.Secondly,in order to implement the multidimensional analysis of health examination data and provide the ability of analyzing key performance indicators from several perspectives to the doctors and managers,we introduce the OLAP technology.The development of multidimensional analysis reports is simplified by using MSTR.The ROLAP server of MSTR can read the fact tables and dimension tables in relational data warehouse,and then transform the related data tables into a unified multidimensional model.We can customize the aggregation results of virtual cubes in the multidimensional model by configuring MSTR,and then provide health examination data multidimensional analysis report services for the doctors and managers.Finally,we discuss the method of health risk appraisal.Through applying classification technology in data mining technology,the relationships between users’medical test results and check conclusions are explored,and the risk forecasting models are established.We select the decision tree、naive Bayes and the support vector machine model,which are three common classification model,to perform experiments on a real health examination data set,and all the accuracies of these three classifiers exceeds 80 percent,which proves that the classification is a feasible method for health risk assessment.In addition,we discuss the imbalance data set problem in the experiment,and choose the over-sampling method to balance the training data.In the comparison experiment,we use SMOTE algorithm to preprocess the training data.The classification abilities of the few classes which we pay more attention improve significantly for all the three classifiers,meanwhile the classification abilities of the majority classes change little,which proves that over-sampling method is a feasible preprocessing method for health examination data set.

  • 【网络出版投稿人】 东北大学
  • 【网络出版年期】2019年 01期
  • 【分类号】TP311.13
  • 【下载频次】104
节点文献中: