节点文献
基于多模态特征融合的遥感图像分类研究
Remote Sensing Image Classification Research Based on Multimodal Feature Fusion
【作者】 朱晨;
【作者基本信息】 西安电子科技大学 , 电子信息硕士(专业学位), 2024, 硕士
【摘要】 遥感图像分类是精准识别地物类型的重要方法,助力环境监测、城市规划、资源调查等领域决策。因为单一模态遥感数据的信息量有限,传统遥感图像分类方法难以应对复杂多变的遥感场景。本文在卷积神经网络、Transformer和自监督学习等深度学习技术的基础上,设计了多模态特征融合的遥感图像分类方法,进一步提取各模态数据特征,实现模态间特征交互融合,研究了样本受限情况下对比学习方法完成分类任务。具体研究工作如下:(1)针对单模态数据信息量受限、卷积神经网络对光谱特征提取不充分难题,结合高光谱数据和激光雷达数据设计了一种基于改进卷积神经网络的多模态遥感图像分类方法。首先设计了双分支空谱残差网络对高光谱数据的空间和光谱特征进行提取;其次在分类任务中增加了激光雷达模态数据,并进行特征提取;最后在特征融合阶段利用孪生神经网络和可学习参数,使得特征融合时可以自适应动态调整。由于增加了激光雷达数据的高程信息和空间特征信息,光谱分支进一步提取了高光谱数据的光谱特征,孪生神经网络和可学习参数使得特征融合更合理,所提出的方法具有更高的分类性能。在Houston2013和Trento数据集上进行了实验,结果表明:在两个数据集上,所提出的方法的分类准确率对比近几年的五个优秀分类方法中最好的提升了1.01%。(2)针对多模态遥感数据的异构特征间难以保证一致性、模态间特征交互不充分的问题,设计了一种双向交互融合的特征融合网络。首先为了更好的提取高维度、信息复杂的高光谱图像特征引入Spectral Former;其次在预训练阶段加入一致性损失,通过反向传播一致性信息保证两个模态间特征的一致性;最后在特征融合阶段加入交叉注意力,促使两个模态特征充分交互,提高分类的准确率,增加模型的鲁棒性。在Houston2013和Trento数据集上进行了实验,结果表明:所提出的方法在各类别中有更好的平衡性,在Houston2013数据集上有6个类别的分类准确率达到最高,分类准确率较除研究内容一外最好的方法提升了1.79%。(3)针对监督方法过于依赖样本数量和质量,而多模态遥感数据样本受限、标注成本高的问题,设计了一种基于自监督对比学习的多模态遥感图像分类方法,利用对比学习进行预训练从无标签样本中学习特征表示。首先利用数据增强构建大量正负样本对帮助提高模型泛化性,使模型更容易区分不同类别样本;其次设置模态内对比学习和模态间对比学习,提升模型模态内特征不变性的同时捕获多模态数据间的而对应关系。最后延续使用上一个研究内容中效果较好的一致性损失和交叉注意力模块辅助特征融合。在Houston2013和Trento数据集上进行了实验,结果表明:所提出的方法的准确率能够在每类仅有10个微调样本的时候远超监督学习的方法,在分类准确率上领先七个对比实验中效果最好的方法2.63%。
【Abstract】 Remote sensing image classification is an important method for accurately identifying feature types,which helps decision-making in the fields of environmental monitoring,urban planning,and resource investigation.Because of the limited amount of information in single-modal remote sensing data,traditional remote sensing image classification methods are difficult to cope with complex and changing remote sensing scenes.In this thesis,on the basis of deep learning techniques such as convolutional neural network(CNN),Transformer and self-supervised learning,we design a remote sensing image classification method for multimodal feature fusion,further extract the features of each modal data,realize the interactive fusion of inter-modal features,and study the comparison of the learning methods to complete the classification task in the case of sample limitation.The specific research works are as follows:(1)Aiming at the problem of limited information of unimodal data and insufficient extraction of spectral features by CNN,a multimodal remote sensing image classification method based on improved CNN is designed by combining hyperspectral data and Li DAR data.Firstly,a two-branch null spectral residual network is designed for spatial and spectral feature extraction of hyperspectral data respectively.Secondly,Li DAR modal data is added to the classification task and feature extraction is performed.Finally,Siamese neural network(SNN)and learnable factors are used in the feature fusion stage,so that the feature fusion can be adjusted adaptively and dynamically.The proposed method has higher classification performance due to the addition of elevation information and spatial feature information of the Li DAR data,the spectral branch further extracts the spectral features of the hyperspectral data,and the SNN and learnable factors make the feature fusion more reasonable.Experiments are carried out on Houston2013 and Trento datasets,and the results show that on both datasets,the classification accuracy of the proposed method improves by 1.01%over the best of the five excellent classification methods in recent years.(2)Aiming at the problems of difficulty in ensuring consistency between heterogeneous features and insufficient inter-modal feature interaction in multimodal remote sensing data,a two-way interactive fusion feature fusion network is designed.Firstly,Spectral-Former is introduced in order to better extract high dimensional and complex hyperspectral image features.Secondly,consistency loss is added in the pre-training stage to ensure the consistency of the features between the two modalities by back propagating the consistency information.Finally,cross-attention is added in the feature fusion stage to promote the full interaction of the two modal features to improve the accuracy of the classification and increase the robustness of the model.Experiments are conducted on Houston2013 and Trento datasets,and the results show that the proposed method has a better balance among the categories and achieves the highest classification accuracy in six categories on the Houston2013 dataset,with an improvement of 1.79%in classification accuracy over the best method except for research element one.(3)Aiming at the problem that supervised methods rely too much on the number and quality of samples,while the samples of multimodal remote sensing data are limited and the labelling cost is high,a multimodal remote sensing image classification method based on self-supervised contrast learning is designed to learn the feature representations from unlabeled samples by using contrast learning for pre-training.Firstly,a large number of positive and negative sample pairs are constructed using data augmentation to help improve the model generalization and make the model easier to distinguish between different categories of samples.Secondly,intra-modal contrast learning and inter-modal contrast learning are set up to improve the feature invariance of the model within the modality while capturing the correspondence between the multi-modal data.Finally,we continue to use the loss of coherence and cross-attention modules,which are more effective in the previous study,to assist feature fusion.Experiments are conducted on the Houston2013 and Trento datasets,and the results show that the accuracy of the proposed method is able to far outperform the supervised learning method when there are only 10 fine-tuned samples per class,and is 2.63%ahead of the best of the seven contrastive trials in terms of classification accuracy.
【Key words】 Remote Sensing Image Classification; Multimodal; Transformer; Self-Attention; Convolutional Neural Network; Contrastive Learning;
- 【网络出版投稿人】 西安电子科技大学 【网络出版年期】2025年 09期
- 【分类号】TP751;TP18