节点文献

基于基类信息辅助的元学习方法研究

Researches on Meta-Learning Methods Based on Base Class Information Boosting

【作者】 刘国栋;

【导师】 何琨;

【作者基本信息】 华中科技大学 , 计算机应用技术, 2023, 硕士

【摘要】 随着深度学习的发展,图像分类任务已经取得巨大成功。然而,深度学习模型的训练往往需要依靠大量数据,当数据较少时由于过拟合会严重损害模型性能,这便产生了小样本问题:当遇到新任务时如何在样本数据较少的情况下仍使模型保持较好性能?元学习是最常用于解决小样本问题的框架,近年来出现了带有预训练的元学习算法,代表工作为元基线算法(Meta-Baseline),其训练过程分为预训练和元训练。然而,随着元训练的进行,元基线算法的准确率会先升后降,最后甚至会弱于原型网络算法,表明预训练得到的表征信息并不能很好地被元训练度量。另外,作为非端到端的算法,元基线算法的元训练必须要等预训练结束后才能进行,且第一二阶段准确率不成正比,意味着必须对多个预训练模型进行元训练。为了解决上述问题,受随机方差消减梯度下降算法(Stochastic Variance Reduced Gradient,SVRG)的启发,设计和实现了一种端到端的基于基类信息辅助的元学习训练方法(Boosting Meta-Training with Base Class Information,Boost-MT),可以更好地利用基类数据信息以辅助元训练过程,从而提升模型性能,同时避免了非端到端训练带来的问题。Boost-MT算法将训练过程分为内外循环两部分。在外循环中,在基类批数据上利用通用机器学习方法计算基类数据外循环损失,但不更新特征提取器的参数,解决了通用机器学习训练得到的特征提取器提取的表征不能被元学习度量的问题。在内循环中,通过元学习计算任务内循环损失,并使用任务外循环损失与基类数据外循环损失共同指导内循环更新,更好地利用了基类信息的类迁移能力。利用常用的小样本数据集对Boost-MT算法进行了实验验证。实验结果表明,随着训练的进行,模型可以快速收敛到稳定的状态,且并未出现性能下降的现象,说明Boost-MT算法更好地利用了基类信息指导元训练。同时,在普通小样本分类任务及跨域任务下,Boost-MT算法的准确率都超过了元基线算法,表明相对于带有预训练阶段的元学习训练方法,所提出的Boost-MT算法更适合执行元学习。

【Abstract】 Deep learning has achieved remarkable success in image classification in recent years.Training these deep learning models often requires a huge amount of data.However,when the available data is limited,the models often perform poorly due to overfitting.This leads to the few-shot learning problem: how to train a model to achieve good results using a small amount of data when faced with new tasks? Meta-learning has recently been the most common framework to address the few-shot learning problem.Meta-learning with pretrained models has emerged in recent years,including the pre-training and meta-training stages.The representative work of this category is Meta-Baseline.However,in the second stage of the Meta-Baseline,the performance on the validation set first increases then decreases,and finally,is even lower than that of the Prototypical Networks,indicating the difficulty of utilizing the representation information of pre-training in the second stage.Besides,Meta-Baseline is not an end-to-end training method,which means the second training stage must wait until the first stage completes convergence,resulting in a significantly longer overall training time.Moreover,because the accuracy of the first and second stages is not proportional,multiple models must be trained in the second stage to achieve better results.To address the above issues,inspired by Stochastic Variance Reduced Gradient(SVRG),a new end-to-end training approach is designed to boost meta-training with base class information,termed Boost-MT,which makes better use of base class information and avoids negative impact in methods which are not end-to-end.In this method,two loops are executed alternately.In the outer loop,Boost-MT calculates the classification loss of one large batch from the base class set but does not update the feature extractor,solving the problem that the representations extracted by the feature extractor trained by general machine learning cannot be measured by meta-learning.In the inner loop,Boost-MT calculates the loss of several divided episodic tasks by meta-learning methods and updates the model by incorporating inner loss and outer loss.Many experiments are conducted to evaluate Boost-MT under some standard few-shot datasets.The results show that the model can quickly converge as the training progresses.No performance degradation indicates that the base class information is better used to guide the meta-training.Boost-MT also outperforms Meta-Baseline in both few-shot classification tasks and few-shot cross-domain tasks,indicating that compared with the meta-learning algorithm with a pre-trained version,the Boost-MT training method is more suitable for performing few-shot learning.

  • 【分类号】TP181
节点文献中: 

本文链接的文献网络图示:

本文的引文网络