节点文献

面向信用风险预测的特征工程研究

Research on Feature Engineering for Credit Risk Prediction

【作者】 李勇;

【导师】 马寿峰;

【作者基本信息】 天津大学 , 系统工程, 2019, 硕士

【摘要】 随着全球数字化货币转型趋势的到来,以及互联网技术与金融行业的深度结合,数字化风控的趋势越来越明显。面对“大数据时代”所积累的海量数据,运用大数据理论分析处理海量数据也对各行各业提出了更高的要求和挑战。信用风险管理控制是金融行业的核心。同时,随着人工智能技术的兴起,金融机构如何运用机器学习、特征工程等技术全面的刻画业务场景、更精准的进行风险预测也成为了当前研究的热点问题。信用风险预测问题本质上是分类问题,结合金融信贷行业数据大多为结构化数据的特点,本文通过对已有文献和方法进行梳理,采用特征工程、机器学习等技术进行研究试验,具体研究贡献总结为以下2方面:1.面向信用风险预测问题提出了具体的特征工程构造方法和流程。首先,针对结构化数据,将数据表抽象为特征实体,提出了对于单个实体聚合和扩展进行特征构造的基本方法。其次,阐述了不同实体之间的连接方式,提出了从基础特征、聚合特征、转换特征、时序特征、组合特征、业务特征六个维度进行特征转换的方法和具体操作。最后运用企业真实结构化数据进行特征工程流程实践,在迭代过程中生成了大量稳定性强,效果好的特征,对于信用风险预测提供了一定的参考和借鉴。2.对比分析特征工程在不同模型上的效果提升,以及对特征重要性进行可解释性分析。分别选择逻辑斯谛回归、支持向量机、随机森林、梯度提升树分类器与交叉验证结合的方式进行训练,并采用多种评估指标综合评价特征效果。结果表明,相对于原始特征,特征工程流程所构造的特征在不同的模型上都有较高的提升效果,特征工程具有实际应用价值。同时,进行特征选择前后对比分析,在训练过程前加入嵌入式特征选择算法进行预训练,筛除无关冗余特征,使得模型效果进一步提升。

【Abstract】 With the advent of the global digital currency transformation trend and the deep integration of Internet technology and financial industry,the trend of digital wind control is becoming more and more obvious.Faced with the massive data accumulated in the "big data era",the use of big data thinking to analyze and process large amounts of data has also raised higher requirements and challenges for all walks of life.Credit risk management control is the core of the financial industry.At the same time,with the rise of artificial intelligence technology,how financial institutions use machine learning,feature engineering and other technologies to comprehensively portray business scenarios and conduct more accurate risk prediction has become a hot topic.problem.The credit risk prediction problem is essentially a classification problem.Combined with the characteristics of financial credit industry data,most of them are structured data.Thesis sorts out existing literature and methods,and adopts feature engineering,machine learning and other technologies.Research trials,the specific research contributions are summarized as the following two points:1.Specific feature engineering construction methods and processes are proposed for credit risk prediction.The specific logic is:for structured data,abstract the data table into feature entities,and propose the basic method of feature construction for single entity aggregation and extension.At the same time,it also expounds the connection between different entities,and proposes the basic features,aggregation features,transformation features,time series features,combination features,business features six dimensions for feature construction methods and specific operations,and finally use the enterprise’s real structured data for feature engineering process practice,generating a large number of stable in the iterative process The characteristics of good results provide a certain reference and reference for credit risk prediction.2.Contrast analysis of the improvement of the effect of feature engineering on different models,and the interpretation of the importance of features.The combination of logistic regression,support vector machine,random forest,gradient boosted tree classifier,and cross-validation was selected for training,and a variety of evaluation indicators were used to comprehensively evaluate the characteristic effects.The results show that compared with the original features,the feature engineering process has a higher improvement effect on different models,and feature engineering has practical application value.At the same time,before and after feature selection comparison analysis,an embedded feature selection algorithm is added before the training process for pre-training to filter out irrelevant redundant features,which further improves the model effect.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2022年 01期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络