节点文献

Boosting族算法在信用评分卡模型中的应用研究

The Application of Boosting Family Algorithm in Credit Rating Card Model

【作者】 王睿;

【导师】 牛一;

【作者基本信息】 大连理工大学 , 应用统计(专业学位), 2021, 硕士

【摘要】 在大数据时代背景中,不断更新和沉淀了巨大量级的交互数据信息,为互联网金融行业带来更全面的参考信息维度,同时也提出了新的挑战。信用评分体系是金融领域中一项里程碑式的成就,它伴随着信用交易成为商业闭环中推崇的新型促销模式这一趋势下,为银行体系、金融机构和消费贷款等公司提供决策标准。信用评分卡作为金融行业的风险控制中的有力抓手,早年便通过Logistics回归方法对其进行模型学习,发展至今已经很是成熟。在此基础上如何将有效信息最大化,利用最新的数据挖掘技术改进对用户的信用评估,具有一定现实意义。本文基于Boosting族算法对模型鲁棒性和准确性提升的优异表现,探究其中的三种算法XGBoost、Light GBM、Cat Boost在信用评分模型上表现,结合Logistics算法对比其在评估指标KS和AUC上是否有提升以及提升程度。第一章对选题基础情况和重点问题进行介绍;第二章为发展历程与研究动态,第一部分从评分卡体系方向对评分体系和评分卡的发展历程,以及具体案例进行介绍。第二部分对国内外学者在评分卡建模的研究情况进行综述;第三章介绍信用评分卡中的基本概念定义和分类情况;第四章从算法原理、算法特点、建模参数三个方面对Boosting族算法介绍;第五部分使用两个真实数据集进行实证分析;第六部分为总结和展望。在实证分析中,对根据两个数据集的自身特点,进行数据清理、特征选取后使用XGBoost、Light GBM、Cat Boost算法分别对两数据集建立单一模型,后使用Blending建立融合模型,发现Boosting族算法建立的单一模型整体上较Logistics有一定的提升,而融合模型并非优于全部单一模型,其中Light GBM算法在第一个数据集中表现最优,Blending融合模型在第二个算法中表现最优。同时,本文使用SMOTE算法对不平衡数据集进行过采样处理得到新数据集,对比原数据发现过采样对单一模型不明显,甚至部分模型评估指标下降;并探究Boosting族算法中的三个单一模型不同调参方式差异,对比其在第一个数据集的原数据和过采样数据结果,在选择不同参数的基础上,发现贝叶斯调参在测试集中AUC值最优。

【Abstract】 In the background of big data era,the huge amount of interactive data information is constantly updated and precipitated,which brings a more comprehensive reference information dimension for the Internet financial industry,and also puts forward new challenges.Credit scoring system is a milestone achievement in the financial field.Along with the trend that credit transaction has become a new promotion mode in the commercial closed-loop,it provides decision-making standards for the banking system,financial institutions and consumer loan companies.Credit scoring card as a powerful grasp in the risk control of the financial industry,early through the logistics regression model learning,development has been very mature.On this basis,how to maximize the effective information,through the latest data mining technology to improve the user’s credit evaluation,has a certain practical significance.Based on the excellent performance of boosting algorithms in improving the robustness and accuracy of the model,this paper explores the performance of the three algorithms xgboost,lightgbm and catboost in the credit scoring model,and compares their improvement in the evaluation index KS and AUC and the degree of improvement by combining with the logistics algorithm.The first chapter introduces the basic situation and key issues of the topic selection;The second chapter is the development process and research trends.The first part introduces the development process of the scoring system and the scoring card from the direction of the scoring card system,as well as specific cases.In the second part,the domestic and foreign scholars’ research on score card modeling is reviewed;The third chapter introduces the basic concept definition and classification of credit score card;The fourth chapter introduces boosting algorithm from three aspects: algorithm principle,algorithm characteristics and modeling parameters;The fifth part uses two real data sets for empirical analysis;The sixth part is the summary and prospect.In the first mock exam,according to the first mock exam of two data sets,data cleaning and feature selection are used.After using XGBoost,Light GBM and Cat Boost algorithms,a single model is set up for two data sets.After that,Blending is used to build the fusion model.It is found that the single model established by Boosting algorithm has a certain improvement over Logistics.The fusion model is the first mock exam.The Light GBM algorithm is the best in the first dataset.The Blending fusion model performs best in the second algorithms.In the first mock exam,we use SMOTE algorithm to process the imbalanced data sets and get new data sets.Compared with the original data,we find that the oversampling is not obvious to the single model,or even the evaluation index of some models is decreasing.The first mock exam and the oversampled data are compared with the results of the three single models in the Boosting family algorithm.On the basis of selecting different parameters,it is found that the AUC value of Bayesian parameter adjustment is the best in the test set.

  • 【分类号】TP311.13;F830
  • 【被引频次】3
  • 【下载频次】470
节点文献中: 

本文链接的文献网络图示:

本文的引文网络