节点文献

基于随机欠采样算法的信用风险研究

Research on Credit Risk Based on Random Under-sampling Algorithm

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 肖衡李莉莉

【Author】 XIAO Heng;LI Li-li;School of Economics,Qingdao University;

【通讯作者】 李莉莉;

【机构】 青岛大学经济学院

【摘要】 针对数据不平衡导致的信用风险识别精度低的问题,利用随机欠采样算法对数据集平衡处理后,采用Logistic回归模型以及随机森林、决策树、XGboost和支持向量机等分类算法分别建立模型并进行预测。实证结果表明,随机欠采样算法可以将信用卡欺诈风险的预测精度从低于75%提升至85%以上,且G-mean和AUC等衡量非平衡数据分类性能的指标均有明显提高,该算法能够有效缓解数据不平衡导致的风险预测性能低下的问题。

【Abstract】 Aiming at the problem of low accuracy of credit risk identification caused by imbalanced data, the random under-sampling algorithm was used to balance the dataset. Then the Logistic regression model and random forest, decision tree, XGboost as well as support vector machine classification algorithms were used to establish the model and make prediction respectively. The research results show that through the random under-sampling algorithm, the prediction accuracy of credit card fraud risk is improved from less than 75% to more than 85%, and the indicators such as G-mean and AUC to measure the performance of imbalanced classification are significantly improved, which can effectively alleviate the problem of low performance of risk prediction caused by data imbalance.

【基金】 国家社会科学基金(批准号:2019BTJ028)资助;山东省金融应用重点研究项目(批准号:2020-JRZZ-03)资助
  • 【文献出处】 青岛大学学报(自然科学版) ,Journal of Qingdao University(Natural Science Edition) , 编辑部邮箱 ,2022年04期
  • 【分类号】TP181
  • 【下载频次】8
节点文献中: 

本文链接的文献网络图示:

本文的引文网络