节点文献

面向小样本数值表数据的自动机器学习优化方法研究

Research on Optimization Method of Automatic Machine Learning for Small Sample Numerical Tabular Data

【作者】 李壮;

【导师】 覃京燕;

【作者基本信息】 北京科技大学 , 计算机科学与技术, 2022, 博士

【摘要】 随着工业革命浪潮的推进,以机器学习为代表的新一代信息技术在国家重点战略领域中发挥着重要作用,但机器学习技术应用在如新材料、物理化学、生物医疗、国防以及各类生产等重点领域往往面临着小样本数值表数据学习的问题,且机器学习技术的成功应用都离不开领域专家的参与,门槛较高且难以推广至其他领域。现有自动机器学习技术虽然一定程度上能够降低机器学习技术在各个领域中的应用门槛,但未针对小样本数值表数据学习的问题进行数据层面的增强,在小样本数值数据问题上仍然面临着泛化性能不足的问题。而且针对模型的评估仅采用建模性能指标,在小样本数值表数据建模上存在过拟合的风险。另外,对于自动机器学习中采用的空间搜索优化方法也存在初始化参数多的问题。针对上述问题,本文以提高面向小样本数值表数据的自动机器学习建模性能为目的,对数据自动增强、数据特征自动构造、参数搜索优化、模型性能评估等方面进行了系统且深入的研究。本文的主要研究内容和创新点如下:1)基于泛化误差界的数据自动持续增强方法:针对小样本数值表数据的自动持续生成筛选问题,研究基于Rademacher复杂度的模型泛化误差界的数据有效性判断,一定程度上避免噪声样本的引入,同时结合域范围扩展的数据生成方法实现了小样本数值表数据的有效持续增强。2)基于指标一致性的数据特征自动构造方法:基于数据分布距离理论,研究影响数据分布效果的主要因素并设计了非重叠度指标,实验验证了非重叠度指标与机器学习建模预测准确度具有较高的一致性。然后基于指标一致性原则分别面向数据平衡和不平衡场景提出了自动特征构造方法GP-ANO和GP-AANO。(1)面向数据平衡的场景,将非重叠度指标引入基于遗传编程的自动特征构造方法,实现了具有良好泛化性能和数据分布效果特征的自动构造,验证了基于指标一致性设计适应度评估的可行性。(2)面向数据不平衡的场景,分析了非重叠度指标的问题并提出增强非重叠度指标,利用增强非重叠度指标,结合建模预测的AUC指标提高了自动特征构造方法在数据不平衡场景的泛化性能。3)基于教与学优化算法的高效优化方法:针对少初始参数的教与学优化算法,研究分析了算法存在的搜索偏见问题,并针对性的引入自适应学习因子,消除搜索偏见。同时针对教与学优化算法探索能力不足,容易陷入局部最优的问题,引入随机自学和变异阶段,增加群体的多样性,有效地提高了算法的优化能力。4)面向小样本的自动机器学习框架:提出一种面向小样本数值表数据的全流程自动机器学习框架,将数据自动增强、特征自动构造、算法参数自动优化以及集成学习融合一体,能够实现数据、特征和算法的全流程增强,并深入分析了集成学习的模型误差与非重叠度指标,设计了基于集成学习误差和改进非重叠度指标的评估指标,然后在材料和医疗领域的实际数据上验证了本文提出的全流程自动机器学习框架的在小样本数值表数据上的有效性。因此,本文研究的方法可以有效地应用于小样本数值表数据的机器学习建模分析场景,且能实现机器学习模型的自动化构建。总的来说,本文的研究和提出的方法对于小样本学习和自动机器学习具有重要的理论研究和实际应用价值。

【Abstract】 With the development of the industrial revolution,the new generation of information technology represented by machine learning has played an important role in the national key strategic areas.However,the small sample numerical tabular data learning problem is encountered when machine learning is applied to the national key strategic areas such as materials science,physical and chemistry,biomedicine,national defense and manufacturing industry.And the successful applications of machine learning depend heavily on the domain experts.Therefore,the threshold of applying machine learning is high and the successful applications are difficult to extend into other areas.Although the existing automatic machine learning technology can reduce the difficulties of applying machine learning into various fields to some extent,the enhancement from the data level is not achieved for small sample numerical tabular data learning problems.The insufficient generalization performance problem is still encountered for the small sample numerical tabular data.Besides,the risk of overfitting is still high with the performance only evaluated by the modeling performance on small sample numerical tabular data.And some optimization methods used in automatic machine learning have too many initialization parameters to be predefined.To improve the performance of automatic machine learning modeling for small sample numerical tabular data,the fields of automatic data enhancement,automatic feature construction,parameter optimization methods,model performance evaluation are systematically and deeply studied in this paper.The main research contents and innovations of this paper are as follows:1)Automatic lasting data enhancement method based on generalization error bounds:Aiming at the problem of automatic lasting data generation and screening of the small sample numerical tabular data,the generalization error bounds based on the Rademacher complexity is studied and adopted to avoid the introduction of noise samples to some extent.Meanwhile,the effective lasting data enhancement of small sample numerical tabular data is achieved by combining with the data generation method of domain extension.2)Automatic data feature construction method based on the consistency of evaluation indexes:Based on the data distribution distance theory,the main factors that affect the data distribution performance are studied and the non-overlap degree index is designed.The experiment has shown that the non-overlap degree is highly consistent with the prediction accuracy of the machine learning model.Then the automatic feature construction methods GP-ANO and GP-AANO for balanced data and imbalanced data are proposed respectively based on the consistent indexes.(1)For the balanced data,the non-overlap degree is introduced into the automatic feature construction method based on genetic programming and the experimental results demonstrated that the automatic feature construction method with the non-overlap degree can achieve a better generalization performance and data distribution.(2)For the imbalanced data,the problem of non-overlap degree is analyzed and the augmented non-overlap degree is proposed.The augmented non-overlap degree combined with the AUC index is used to improve the generalization performance of the automatic feature construction method for the imbalanced data.3)Efficient optimization method based on teaching-learning based optimization algorithm:Teaching-learning based optimization algorithm has an advantage of needing fewer initial parameters.While,it has a search bias to the origin.The search bias is firstly analyzed,and then the adaptive learning factors are introduced to eliminate the search bias.At the same time,the random self-learning and mutation stages are introduced to increase the diversity of the population,preventing trapping into local optimal and improving the optimization performance.4)Automatic machine learning framework for small sample data:An automatic machine learning framework for the small sample numerical data is proposed based on the fusion of automatic data enhancement,automatic feature construction,automatic algorithm parameter optimization and ensemble learning.The error of ensemble learning and the overlap degree are deeply analyzed and an evaluation index based on the ensemble learning error and the improved augmented non-overlap degree is designed.The proposed automatic machine learning framework is validated on the practical data modeling problems of materials and biomedical fields.Therefore,the method studied in this thesis can be effectively applied to small sample numerical tabular data learning problem and the optimal model can be obtained automatically.Overall,the researches and proposed methods in this thesis have important theoretical and practical application value for same sample learning and automatic machine learning.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络