节点文献

面向信贷违约风险的不平衡数据分类和特征选择

Imbalanced Data Classification and Feature Selection for Credit Default Risk

【作者】 韩旭

【导师】 马寿峰;

【作者基本信息】 天津大学 , 系统工程, 2019, 硕士

【摘要】 近些年来,随着国民经济的持续增长,信贷业务开始出现在人们的日常生活之中。虽然信用贷款给我们带来了很多的便利,但是申请用户的与日俱增也给这个行业以及管理者的正常运营带来了巨大挑战。与此同时,信贷数据还存在着类别比例不平衡且数据特征维数过多的问题。在此背景下,如何准确地预测用户是否会存在违约的情况成为了一项亟待解决的难题。本文针对信贷数据存在的不平衡且数据特征维数过高的特点,整理国内外信贷违约风险领域的相关研究文献,同时运用机器学习算法进行建模分析,较为创新性地完成了以下两项工作:一、为了提高分类算法在信贷不平衡数据的预测性能,提出了一种基于高斯混合模型(Gaussian mixture model,GMM)和SMOTE(Synthetic Minority Oversampling Technique)算法相结合的组合采样算法(GSRA),并将GSRA算法应用在信贷数据的不平衡特性之中。GSRA算法利用了高斯混合模型对多数类样本进行欠采样(Under sampling),同时对少数类样本进行SMOTE技术过采样(Over sampling),从而消除了样本间的类别不平衡问题。实验中将GSRA算法与十种常用的分析不平衡数据集的重采样方法以及在有无重采样方法下的性能进行了对比,并且进行了算法的鲁棒性分析。从实验结论中可以看出,GSRA算法明显地增强了分类器的学习性能。同时,该算法对信贷数据集具有较强的鲁棒性与抗噪性。二、针对信贷风险领域数据集存在数据维数过高的问题,设计了一种基于克隆选择原理的特征选择算法(CFSA)。该算法利用生物界的克隆选择原理来指导特征选择过程,并利用现实世界的汽车金融企业数据,进行建模分析。实验研究了有无特征选择算法的分类性能以及CFSA算法与遗传算法(Genetic algorithm,GA)方法的比较,结果证明,特征选择可以有效删除冗余特征提升分类器性能,同时CFSA算法在四个传统分类器中均有更佳的表现结果。此外,实验也验证了经过该算法选择后的特征在现实生活中是具有一定的实际意义的。

【Abstract】 Recently,as China’s social economy keeps growing,credit business has begun to appear in people’s daily lives.Although credit loans have brought us a lot of convenience,the increasing number of applicants has brought enormous challenges to the normal operation of this industry and managers.At the same time,credit data still has the problem of imbalanced class proportion and excessive data feature dimension.In this context,how to quickly and accurately predict whether users will default has become an urgent problem to be solved.In view of the imbalance of credit data and the high dimension of data characteristics,this paper collates the relevant research literature in the field of credit default risk at home and abroad,and uses machine learning algorithm to model and analyze,and accomplishes the following two tasks innovatively:Firstly,in order to improve the prediction performance of classification algorithm in imbalanced credit data,a combined sampling algorithm(GSRA)based on the combination of the Gauss mixture model(GMM)and the Synthetic Minority Oversampling Technology(SMOTE)algorithm is proposed,and the GSRA algorithm is applied to the imbalanced characteristics of credit data.The GSRA algorithm uses the Gauss mixture model to undersampling most samples and SMOTE over sampling for a few samples,thus eliminating the class imbalance between samples.In the experiment,the performance of GSRA algorithm is compared with that of ten common resampling methods for analyzing imbalanced data sets and with or without resampling methods,and the robustness of the algorithm is analyzed.From the conclusion of the experiment,it can be seen that the GSRA algorithm can effectively enhance the learning performance of the classifier.At the same time,the algorithm has strong robustness and anti-noise for credit data sets.Secondly,aiming at the problem of high data dimension in credit data set,a feature selection algorithm based on clonal selection and cohesion index(CFSA)is proposed.The algorithm uses the clonal selection mechanism of biology to guide the feature selection process,and uses the real world data of automobile finance enterprises to model and analyze.Experiments show that feature selection can effectively remove redundant features and improve the performance of classifiers.At the same time,CFSA algorithm has better performance in four traditional classifiers.In addition,the experiments also verify that the features selected by the algorithm have a certain practical significance in real life.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2021年 06期
  • 【分类号】TP311.13;F832.4
  • 【被引频次】1
  • 【下载频次】117
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络