节点文献

监所在押人员风险等级类别的预测研究

Research on Forecasting the Risk Classification of Prisoners

【作者】 王悦

【导师】 冶继民; 潘谷;

【作者基本信息】 西安电子科技大学 , 应用统计硕士(专业学位), 2021, 硕士

【摘要】 监所是对所内在押人员进行教育改造的地方,然而近年来越狱逃脱、自杀行凶等危险性事件在国内外频繁发生。这给监所和社会带来巨大损失的同时,还可能对人民生命安全和国家安全产生威胁。为了推动监所的规范和安全建设,对在押人员的危险管理成为监所建设的重点问题。如何正确利用监所的海量数据并挖掘出有效信息,成为了重中之重的核心问题。目前各国学者在监所领域的各种问题研究中,基本采用机器学习或深度学习的技术从大量数据信息中挖掘潜在规律,从而展开深入研究。基于识别监所风险等级这一研究目标,针对监所数据的不平衡特性这一关键难点,从数据采样和算法两个方面改进,分别提出了基于过采样和欠采样的混合采样(MWMOTE-ENN)风险识别算法和基于集成的风险识别改进AdaBoost算法。针对不平衡多类别数据,从数据水平的角度提出了一种基于聚类思想的过采样方法和欠采样方法结合的混合采样方法:MWMOTE-ENN。利用MWMOTE-ENN方法对监所数据进行采样使类别样本量达到相对平衡的状态,解决了少数类分布不均和边界模糊的问题,然后采用AdaBoost算法进行训练并预测风险等级。为验证基于MWMOTE-ENN风险预测算法的有效性,对比了欠采样、过采样、混合采样等数据采样方式,并以召回率、精确度、F值和G均值等指标作为评价标准。实验结果显示,基于MWMOTE-ENN的风险分类算法在整体分类效果和各类别召回率上均优于其他采样算法,将三个有风险等级的识别率提高了20个百分点左右。针对不平衡数据导致的类别损失差异问题,提出了一种修正的AdaBoost算法:R-AdaBoost。通过调整少数类的初始权重和误差函数,以及在分类器权重中引入不平衡因子和误判因子,从而筛选出对少数类友好的分类器,利于降低有风险样本的错分率。为了验证R-AdaBoost算法的效果,采用5折交叉验证,从均衡和不平衡两个角度分别对比了R-AdaBoost算法和AdaBoost算法的分类效果。实验结果表明:(1)两种算法在均衡采样时的分类效果都优于不平衡场景。(2)在均衡场景下,两种算法的整体效果不分伯仲,但R-AdaBoost算法在少数类别的召回率和误判率方面略优。(3)在不平衡场景下,随着不平衡度增加,R-AdaBoost算法的识别能力逐渐优于原有算法。这表明,R-AdaBoost算法有效提高了对风险类别等级的预测能力,降低了三个有风险类别的漏判率,在不平衡情形下的分类效果显著优于原有算法。

【Abstract】 Prison is a place to educate and reform the prisoners.However,in recent years,dangerous events such as prison break and suicide have taken place recurrently at home and abroad.While this brings huge losses to prisons and society,it may also threaten people’s lives and national security.In order to improve the standardized and safety construction of prisons,the risk management of prisoners has become a crucial problem in the construction of prisons.How to make good use of the massive data and mine the effective information has become the most important core problem.At present,in the research of various problems of prisons,scholars have basically adopted machine learning or deep learning technology to dig out potential laws from a large number of information and carry out in-depth research.Based on the research goal of prison risk level identification and the key difficulty of the unbalanced characteristics of prison data,a hybrid sampling risk identification algorithm(MWMOTE-ENN)is proposed from the aspect of data sampling,which is based on oversampling and under-sampling.Besides,an improved risk identification algorithm based on the integration is proposed from the aspect of algorithm improvement.For imbalanced multi-category data,a hybrid sampling method is proposed from the perspective of data level: MWMOTE-ENN,which combines oversampling method based on clustering idea and under-sampling method.The MWMOTE-ENN method is used to sample the prison data,so that the sample size of the categories reaches a relatively balanced state.The problems of uneven distribution of minority classes and fuzzy boundaries are solved.Then the AdaBoost classification algorithm is used to train the data and predict the risk level.In order to verify the effectiveness of the risk prediction algorithm based on MWMOTE-ENN,this thesis also compares with under-sampling,over-sampling and mixed sampling methods.Indexes such as recall rate,precision,F-value and G-mean value are used as the evaluation criteria.The results of our experiments demonstrate that the risk classification algorithm based on MWMOTE-ENN is superior to several other sampling algorithms in overall classification effect and recall rate of each category.The recognition rate of the three risky levels has been increased by about 20 percentage points.Aiming at the problem of class loss differences caused by unbalanced data,a revised AdaBoost algorithm(R-AdaBoost)is proposed.By adjusting the initial weight and error function of the minority class,and introducing the imbalance factor and misjudgment factor into the weight of the classifier,a classifier that is friendly to the minority class can be selected.It is beneficial to reduce the error rate of risky samples.For the sake of the effect of the R-AdaBoost algorithm,this thesis uses 5-fold cross-validation to compare the classification effect of the R-AdaBoost algorithm and the AdaBoost algorithm from two perspectives of equilibrium and imbalance.The experimental results show that:(1)The classification effect of the two algorithms in balanced sampling is better than that of unbalanced scenes.(2)In the balanced scenario,the overall effects of the two algorithms are equal,but the R-AdaBoost algorithm is slightly better in the recall rate and misjudgment rate of minority categories.(3)In the unbalanced scene,with the increase of imbalance degree,the recognition ability of the R-AdaBoost algorithm is better than the original algorithm.This shows that the R-AdaBoost algorithm effectively improves the predictive ability of risk category level,reduces the missed judgment rate of three risky categories.The classification effect of the R-AdaBoost algorithm is significantly better than the original algorithm in the unbalanced situation.

  • 【分类号】D926.7;D631.7;TP18
  • 【下载频次】65
节点文献中: