节点文献
大数据下医保欺诈的有效识别模型
Effective Identification Model of Medical Insurance Fraud for Big Data
【摘要】 针对现在社会医保诈骗问题,提出了大数据下医保欺诈的有效识别模型.首先运用excel对数据进行预处理,建立数据挖掘有效识别数据集;其次通过主成分分析构建欺诈识别的有效指标体系;再次由K-Means聚类得到可疑的医保欺诈行为的类别,并由判别分析中的交叉确认估计来确认可疑行为判断类别的准确性.随后,由因子分析中的数据映射关系找到与欺骗行为有关的科室、医生、医嘱子类,并把欺诈行为归为医疗保险服务供应方的诈骗行为、医疗保险需求方的诈骗行为和医疗保险服务供应方与需求方合谋的诈骗行为这三大类;最后把模型用于由样本经验分布的反函数生成的大数据中,解决了统计分析中样本少而使统计分析出现误差这一问题.
【Abstract】 Aiming at the problem of social health insurance fraud, an effective identification model for Medicare fraud in big data is proposed. First, the excel is used to preprocess the data, establishing an effective identification data mining data set. Secondly, the effective index system of fraud identification by principal component analysis is constructed. The category of suspicious medical fraud is obtained from K-Means clustering, and the accuracy of the suspicious behavior judgment category by cross validation of discriminant analysis is confirmed. From the data mapping relationship of factor analysis is used to find the deceptive behavior of the departments, doctors, medical sub-class. Then, the fraud is classified as medical insurance service provider, demander, collusion between supply and demand three categories. Finally, the model for big data which generated by the inverse of the empirical distribution of the samples, is used to solve the problem of statistical analysis of the error for the sample less.
【Key words】 effective identification data set; principal component analysis; K-Means clustering; discriminant analysis; factor analysis; big data;
- 【文献出处】 汕头大学学报(自然科学版) ,Journal of Shantou University(Natural Science Edition) , 编辑部邮箱 ,2018年01期
- 【分类号】R197.1;TP311.13
- 【被引频次】14
- 【下载频次】702