节点文献

基于表示学习的图像融合算法研究与应用

Research and Application on Representation Learning Based Image Fusion Algorithm

【作者】 李辉;

【导师】 吴小俊;

【作者基本信息】 江南大学 , 控制科学与工程, 2021, 博士

【摘要】 信息融合技术尤其是多模态信息融合技术,在军事国防、民用安防以及智慧城市建设等领域都有着广泛的应用,并时刻影响着人们的日常生活。同时伴随着这些领域的快速发展,也涌现出了大量的全新的视觉数据。从早期的单模态可见光数据(RGB)到使用红外摄像机采集的热红外和近红外数据,再到更加专业的深度相机、高光谱相机(卫星)以及医学设备采集的视觉数据,这些都显示出实际应用场景对多模态信息融合技术的迫切需求。基于这些视觉数据以及实际需求,多模态视觉信息融合始终是一个热门的研究课题。而如何将不同设备采集到的视觉信息进行有效地整合、处理也变得尤为重要。在本文中,多模态视觉信息融合主要是指多模态图像融合,其主要目的是将多幅配准好的多源(多模态)图像进行整合,使得在最终的融合图像中既能够保留不同模态的有效特征又能够减少模态间的冗余信息。最终的融合图像能够包含更多的有用信息,有利于提高下游计算机视觉任务的性能,如目标跟踪、语义分割、显著性检测等。经过近三十多年的发展,研究学者们提出了许多图像特征提取算法、融合策略以及融合模型。在计算机视觉发展的每一个阶段,都出现了大量结合最新技术的图像融合算法。而随着机器学习的飞速发展,在图像融合领域,表示学习技术也受到了研究学者的高度关注。从早期的基于多尺度变换的融合算法,到基于稀疏/低秩表示的图像融合模型,再到结合其他图像处理技术(形态学分析,脉冲耦合神经网络等)的融合算法,都体现出图像融合任务持续的研究热度。而深度学习的出现,又为融合算法的研究注入了新的活力。利用深度学习强大的学习能力,将大量需要人为设定的处理操作用神经网络模块来代替,极大降低了图像特征提取和融合策略设计的难度。具体来说,基于多尺度变换的融合模型是利用频域信号处理技术(小波分解,剪切波变换,轮廓波变换等)将源图像变换到频域,然后使用恰当的融合策略将频域系数进行融合,最后利用多尺度分解的逆变换得到融合图像。而基于稀疏/低秩表示的融合算法则是直接在空域对图像进行操作,先将图像分为多个图像块并将其重新组成一个样本矩阵,然后使用稀疏/低秩表示来计算样本矩阵对应的稀疏/低秩系数,最终使用恰当的融合策略以及重构模型得到融合图像。基于深度学习的融合算法可以分为三类:(1)基于预训练网络;(2)基于自编码网络;(3)基于端对端网络。基于预训练网络的融合算法是利用预训练的大规模网络强大的特征提取能力来提取源图像特征(代替了传统的图像特征提取方法),然后使用恰当的融合策略计算多模态图像之间的融合权重,得到最终的融合图像。而基于自编码网络结构的融合模型,则是将图像特征提取和融合图像生成两个操作用神经网络代替,从而得到更好的融合图像。基于端对端网络的图像融合网络模型是将所有的融合操作都用网络代替,利用网络结构和损失函数的设计来达到期望的融合效果。本文从以上几个方面对多模态图像融合算法进行了探索。首先分析了稀疏表示在图像融合中的不足,对其进行针对性的改进。针对深度学习的特点以及训练图像融合网络的痛点问题,本文也对基于深度学习的融合算法进行了研究与探索。本文的主要工作概括如下:(1)提出基于低秩表示的图像融合算法,这是低秩表示首次被引入图像融合领域。基于字典学习和低秩表示的融合算法利用梯度方向直方图(Histograms of Oriented Gradients,HOG)特征对图像块进行分类,并学习一个全局字典,然后利用全局字典和低秩表示得到融合图像。在基于低秩表示融合算法的基础上,又提出了基于多级图像分解和潜在低秩表示的融合算法,根据潜在低秩表示对图像分解的特点将其推广到深度领域,提出图像的多级分解模型,有效地增强了算法的融合性能。(2)提出基于预训练神经网络的图像融合算法,将大规模神经网络引入图像融合任务中。基于VGG-19网络和多级深度特征的图像融合算法,是将预训练的VGG-19网络用于提取图像的多级深度特征,然后得到对应于源图像的融合权重,最终对原图进行融合。此外,作者又将预训练的Res Net-50网络引入到融合领域,并使用零相位成分分析(Zero-phase Component Analysis,ZCA)对深度特征进行处理,得到更准确的融合权重。这一类算法进一步提升了融合性能,并为图像融合领域提供了全新的探索方向。(3)提出基于自编码网络框架的图像融合算法。这类方法旨在加强深度学习与图像融合任务的有机结合,并探索深度学习与图像融合任务更多的可能性。首先训练一个自编码网络(由编码器和解码器组成),对网络的输入进行重构。编码器用来提取图像深度特征,解码器则利用得到的深度特征对原始输入进行重构。这种训练方式并不需要特定的训练样本,因此能够有效避免图像融合领域训练样本不足的问题。在测试时,针对不同类型的融合图像,引入恰当的融合策略,最终得到融合图像。这类方法展现出了强大的灵活性和可扩展性,受到大量研究学者的关注。(4)提出基于端对端网络的图像融合算法。尽管基于自编码网络的融合算法取得了优异的融合性能,但是这类方法的融合策略还是需要人为设定。针对这一问题,本文对端对端的融合网络进行了探索。第一种方式是用一个可学习的简单网络代替基于自编码网络融合算法中的融合策略,通过与自编码网络框架结合,提升了算法的融合性能。第二种方式是提出全新的网络结构和损失函数,将图像融合的所有操作用一个网络直接实现,输出为融合图像。这类算法不仅提高了融合性能,并且能够很方便的应用于其他计算机视觉任务。

【Abstract】 Information fusion technique(especially multi-modal information fusion technique)has been widely applied into many fields and affects people’s daily life,such as military and national defense,civilian security and construction of smart city.With the rise of these fields,there are a lot of new visual data.From one modality visual information(RGB image)to the infrared(thermal,near)image,even the depth image(depth camera),hyperspectral image(satellite)and medical image(medical equipment),these multi-modal data show the urgent need for information fusion technique in real application scenarios.Thus,the multi-modal information fusion technique still is a popular research topic.It is very important to combine,process and utilize multi-modal information obtained from different vision devices.In this thesis,the multi-modal information fusion mainly refers to the multi-modal image fusion.The main purpose of image fusion is to generate a single composite image which contains more complementary features and reduces redundant information from multi-modal source images.Ideally,the fusion algorithm can preserve more useful information into the final fused image,which is benefit for the down-stream computer vision tasks,such as object tracking,semantic segmentation and salience detection.After 30 years development,in image fusion field,a lot of algorithms are proposed,including feature extraction algorithms,fusion strategies and fusion models.In each milestone of computer vision development,the image fusion task is still attracted much attention.In recent years,with the rise of machine learning,the representation learning technique is attached attention by many researchers.From the traditional image fusion methods(multi-scale transform,sparse/low-rank representation)to the hybrid fusion algorithms(Morphological Analysis,Pulse-coupled neural networks etc.),which reflects the continued research enthusiasm for image fusion task.Moreover,with the learning mechanism(deep learning),it injects new vitality into the image fusion research.And the design complexities of feature extraction and fusion strategy can be reduced.Thus,we can easily design the network architecture for the specific image fusion task.For multi-scale transform based fusion algorithms,they are all based on signal processing techniques(such as wavelet,shearlet,contourlet).Firstly,the aim of these methods is to transform the source data into frequency domain and obtain multi-scale features.Then,with appropriate fusion strategy and the inverse transform,the final fused images are generated.On the contrary,the sparse/low-rank representation(SR/LRR)based fusion models directly process the source data in spatial domain without any lose of information in transform processing.In SR/LRR based methods,the source images are divided into several image patches and regrouped into a new sample matrix.Then,the SR/LRR are utilized to calculate the coefficients of sample matrix.With appropriate fusion strategy for coefficients and the reconstruction model,the final fused images can be obtained.For deep learning based fusion models,they can be categorized into three classes:(1)Pretrained network based framework;(2)Auto-encoder based framework;(3)End-to-end network based fusion framework.In pre-trained network based framework,the pre-trained neural networks which replace the traditional feature extraction processing are utilized to extract the deep features from source images.Then,appropriate fusion strategies are designed to calculate the fusion weights based on the extracted deep features.Finally,the fused images are generated by the weights and the source images.For the auto-encoder based fusion model,the encoder and the decoder are utilized to extract features and generate the fused image.The main difference between the pre-trained network based fusion framework and the auto-encoder based fusion model is that the architecture can be designed for a specific fusion task in auto-encoder based fusion model.In the end-to-end network based fusion framework,all the image fusion processes are replaced by the specific designed fusion network.With the appropriate architecture and the designed loss functions,it can achieve better fusion performance.In this thesis,we explore the image fusion theory from the above directions.Firstly,we analyze the drawbacks of SR in image fusion field and improve the SR based fusion methods which achieves better fusion performance.Furthermore,according to the main drawbacks of deep learning based fusion methods,we propose several schemes to achieve better fusion performance in this direction.The main contributions include:(1)Low-rank representation(LRR)and dictionary learning(DL)based image fusion methods are proposed.To the best of my knowledge,this is the first time that low-rank representation is applied into image fusion field.Firstly,we combine the DL and the LRR to extract the local and global features from source images.The histograms of oriented gradients(HOG)is utilized to classify the image patches which are divided from source images,and a global dictionary can be learned by these patches.With the global dictionary and LRR,the fused image can be obtained.Furthermore,based on the LRR fusion model,a multi-level decomposition and latent low-rank representation(Lat LRR)based fusion algorithm is proposed,in which the Lat LRR is generalized into a deep version and obtains better fusion performance.(2)The pre-trained neural network based fusion methods are proposed.This is also the first time that the pre-trained large networks are applied into image fusion tasks.Due to the insufficient training data in image fusion field,in early stage,it is difficult to train a specific deep neural network for fusion tasks.Thus,we propose a pre-trained VGG-19 and multi-layer deep features based fusion framework.The VGG-19(trained on Image Net) is utilized to extract the multi-level deep features from source images,the fusion weights of each source image are obtained by several appropriate fusion strategies.Finally,the fused images are reconstructed by the fusion weights and source images.Moreover,we also apply the Res Net-50 into image fusion field and the zero-phase component analysis (ZCA)is used to purge the deep features.The fusion performance is improved by above frameworks,and a new research direction of image fusion is provided.(3)The auto-encoder based fusion models are proposed.In these models,we attempt to explore more possibilities between deep learning and image fusion under a limit condition (insufficient training data).Firstly,an auto-encoder network is designed to reconstruct the input.The encoder is used to extract deep features and the decoder is utilized to reconstruct the input.In this training framework,we do not need a specific training data to train the auto-encoder,which can partially solve the main problem in image fusion (insufficient training data).In testing phase,for different fusion tasks,we can apply different fusion strategies to obtain the fused image.These fusion models show powerful flexibility and expansibility,and a lot of researchers pay more and more attention on this direction.(4)The end-to-end network based fusion frameworks are proposed.Although the auto-encoder based fusion models achieve better fusion performance,the fusion strategy is still designed manually.Thanks to the large multi-modal datasets,we can study on the end-to-end fusion network.To reduce the artificial influence in auto-encoder based models,we propose a learnable network to replace the fusion strategy.With the loss functions and the multi-modal training dataset,we obtain better fusion performance and better quality of fused images.Furthermore,a novel LRR based network architecture is proposed and a multi-level loss function is introduced to maintain the salient information into fused images.These models are also easy to combine with other computer vision tasks and improve the performance.

  • 【网络出版投稿人】 江南大学
  • 【网络出版年期】2022年 03期
节点文献中: