节点文献
基于集成学习的车贷违约预测研究
Research on Automotive Loan Default Prediction Based on Ensemble Learning
【作者】 李雪梅;
【导师】 张旭;
【作者基本信息】 大连理工大学 , 应用统计, 2024, 硕士
【摘要】 当前汽车行业发展十分迅速,同时也带动了车贷业务的高速发展。汽车贷款具有手续审批快且免抵押等因素,因此受到广大群众的欢迎。各大汽车贷款平台积极响应国家号召,探索汽车消费信贷领域,不断地创新车贷行业的相关业务,适应并满足广大客户的需求。然而,由于越来越多的人在购买汽车时会采用贷款的方式,车贷的违约率也在逐渐的增长,这会给提供车贷的企业主体带来十分重大的损失,违约客户的激增,导致车贷平台深深陷入了发展危机。因此,利用大数据和数据挖掘技术建立准确、高效的风险评估模型具有很强的现实意义。本文基于上述车贷业务存在的问题,使用国内某贷款机构提供的部分贷款数据进行统计分析,并建立了车贷违约风险预测模型,使得贷款机构能够提高识别车贷风险客户的能力,从而实现风险可控,最终减少公司的经济损失。本次研究主要包括以下几部分:首先对主要特征变量进行描述性统计,对数据有一个初步的认识,并对数据进行预处理提高数据质量;然后用包装法和嵌入法相结合的方法进行特征选择,共筛选出31个对目标变量具有显著性影响的特征;接下来针对数据不平衡的问题,采用SMOTE过采样的方法进行数据平衡化处理;最后进行实证分析,分别建立了随机森林模型、XGBoost模型、Light GBM模型,并采用随机搜索方法优化模型,然后将其作为初级分类器,将逻辑回归模型作为次级分类器,构建Stacking融合模型并对Stacking模型进行改进优化,使用模型评价指标AUC值、F1值、召回率、精确率来对比单一模型和Stacking融合模型的预测性能。实证研究表明,Stacking模型融合后,其预测性能是优于单一分类模型的,其AUC值为0.917,精确率为0.906,召回率为0.784、F1值为0.841;对模型改进后,进行权重分配的Stacking模型其各个评估指标均是最优的,其AUC值为0.931,精确率为0.912,召回率为0.803、F1值为0.854。因此,加权的Stacking融合模型具备良好的预测性能,能够实现对车贷风险客户的准确识别。
【Abstract】 The current rapid development of the automotive industry has also driven the rapid growth of car loan business.Automobile loans are popular among the general public due to their fast approval process and no mortgage requirements.However,as more and more people use loans to purchase cars,the default rate of car loans is gradually increasing,which will bring significant losses to the business entities providing car loans.The surge in default customers has led to a deep development crisis for car loan platforms.Therefore,using big data and data mining techniques to establish accurate and efficient risk assessment models has strong practical significance.This article is based on the problems existing in the car loan business mentioned above,using partial loan data provided by a domestic lending institution for statistical analysis,and establishing a car loan default risk prediction model.This enables lending institutions to improve their ability to identify car loan risk customers,thereby achieving controllable risks and ultimately reducing the company’s economic losses.This study mainly includes the following parts: Firstly,descriptive statistics were conducted on the main characteristic variables to gain a preliminary understanding of the data,and data preprocessing was carried out to improve data quality;Then,a combination of packaging and embedding methods was used for feature selection,and a total of 31 were selected through screening characteristics that have a significant impact on the target variable;Next,to address the issue of data imbalance,the SMOTE oversampling method will be used for data balancing processing;Finally,the empirical analysis was carried out,and the random forest model,XGBoost model,and Light GBM model were established respectively,and the random search method was used to optimize the model.Then,as the primary classifier,the logical regression model was used as the secondary classifier,and the Stacking fusion model was constructed and optimized.The prediction performance of the single model and the Stacking fusion model was compared using the model evaluation indicators AUC value,F1 value,recall rate,and accuracy rate.Empirical studies have shown that the predictive performance of the Stacking model after fusion is superior to that of a single classification model,with an AUC value of 0.917,an accuracy rate of 0.906,a recall rate of 0.784,and an F1 value of 0.841;After improving the model,the Stacking model with weight allocation has the best evaluation indicators,with an AUC value of 0.931,an accuracy rate of 0.912,a recall rate of 0.803,and an F1 value of 0.854.Therefore,the weighted Stacking fusion model has good predictive performance and can accurately identify car loan risk customers.
【Key words】 car loan default prediction; feature selection; SMOTE methord; Stacking model fusion;
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2025年 08期
- 【分类号】F832.4;F426.471;TP18