节点文献
基于注意力的跨模态融合行人重识别技术研究
A Research on Attention-Based Cross-Modal Person Re-identification
【作者】 张旭;
【导师】 陈杰;
【作者基本信息】 北京邮电大学 , 电子科学与技术, 2025, 硕士
【摘要】 行人重识别技术旨在实现跨摄像头行人检索,然而传统方法无法解决夜间的弱光图像带来的效果下降问题。为弥补这一缺陷,基于可见光-红外图像的跨模态行人重识别技术被提出。但红外图像的引入带来巨大模态差异,同时令行人特征对齐难度增加。现有跨模态行人重识别方法一方面依赖模态共享特征缩小模态差异但忽视行人特征特异性,使行人特征判别性下降;另一方传统方法不能解决全局特征错位降低检索精度的问题。针对上述问题,本文主要进行了如下工作。针对不同模态特征差异大、行人特征判别性不足的问题,提出基于全局注意力的多尺度特征融合的跨模态行人重识别算法,使可见光-红外模态特征差异降低、行人特征判别性提高,实现行人检索精度上升。该算法首先设计了双流全局注意力模块,通过通道-空间注意力指导浅层特征中的特异性特征与深层特征中模态共享特征融合;同时引入中心分离损失,使相同行人聚集、不同行人分散,帮助模型在训练过程中建立两种特征的动态平衡。在SYSU-MM01、RegDB数据集进行对比与消融实验,其Rank-1精度分别达到74.5%与90.8%,相较基线算法可以降低16.59%的模态差异,提升25.71%行人相对间距,验证了算法降低模态差异、增强判别性能力。针对行人特征错位、全局特征匹配效果差的问题,提出基于局部注意力的特征对齐的跨模态行人重识别算法,实现行人特征按语义对齐、全局-局部特征协作匹配,令其在多尺度特征融合算法基础上减少错位误差、提升检索精度。首先,算法特征对齐模块提取、拼接具有不同语义局部特征,实现行人特征按语义对齐;而后利用局部特征对齐损失进一步提升相同语义特征相似性;最后基于局部特征的亲和度匹配算法按局部特征块计算相似度并利用亲和度信息修正,实现全局-局部特征协作匹配。通过消融实验验证特征对齐可以降低特征错位误差,实现4%以上的特征距离修正,细粒度匹配可以提升5%以上的mAP精度。将二者结合,组成最终的基于注意力的跨模态融合行人重识别算法,首先利用基于全局注意力的多尺度特征融合算法进行行人特征提取融合,而后将其输出作为基于局部注意力的特征对齐算法的输入,得到语义对齐的混合特征,最后利用基于局部特征的亲和度匹配算法得到最终的检索结果。经实验,基于注意力的跨模态融合行人重识别算法在LLCM、SYSU-MM01和RegDB数据集上进行性能测试,其Rank-1 精度分别达到了 71.0%、77.5%、95.0%。
【Abstract】 Person re-identification aims to achieve cross-camera person retrieval.However,traditional methods cannot address the performance degradation caused by low-light images in nighttime conditions.To overcome this limitation,visible-infrared cross-modal person re-identification has been proposed.The introduction of infrared images,however,introduces significant modality gaps while exacerbating feature misalignment in pedestrian representations.Existing cross-modality re-identification methods attempt to mitigate modality gaps through modality-shared features but overlook modality-specific characteristics,thereby diminishing the discriminative power of pedestrian features.Meanwhile,conventional approaches fail to address the global feature misalignment that compromises retrieval accuracy.To address these challenges,this thesis conducts the following work:To address the challenges of significant modality gaps and insufficient feature discriminability,we propose a cross-modality person reidentification algorithm based on multi-scale feature fusion with global attention,which reduces infrared-visible modality gaps,enhances pedestrian feature discriminability,and improves retrieval accuracy.The algorithm first designs a dual-stream global attention module that employs channel-spatial attention mechanisms to guide the aggregation of modalityspecific features from shallow layers and modality-shared features from deep layers.Simultaneously,a center separation loss is introduced to enforce intra-class compactness and inter-class separation,helping the model establish dynamic balance between these two types of features during training.Comparative and ablation experiments on SYSU-MM01 and RegDB datasets achieve Rank-1 accuracies of 74.5%and 90.8%,respectively.Compared to baseline algorithms,it reduces the modality gaps by 16.59%and increases pedestrian relative spacing by 25.71%,validating its effectiveness in narrowing modality gaps and enhancing discriminability.To address the challenges of pedestrian feature misalignment and the inadequacy of global features for matching,we further propose a crossmodality re-identification algorithm with local attention-based feature alignment.This approach achieves semantic-aligned person features and collaborative global-local feature matching,reducing misalignment errors and improving retrieval accuracy beyond the baseline multi-scale feature fusion method.First,the algorithm constructs a feature alignment module to extract and concatenate semantically distinct local features,aligning pedestrian features according to semantic consistency.Subsequently,a local feature alignment loss further enhances similarity of same-semantic features.Finally,an affinity matching algorithm based on local feature blocks calculates similarity scores and corrects them using affinity information,realizing collaborative global-local matching.Ablation experiments demonstrate that feature alignment reduces feature distance over 4%by reducing feature misalignment,while fine-grained matching improves mAP accuracy by over 5%.Integrating the two components,we form the final attention-based cross-modal person re-identification algorithm.First,the multi-scale feature fusion algorithm based on global attention is utilized to extract and fuse person features.Subsequently,its output serves as the input to the feature alignment algorithm based on local attention,yielding semantically aligned hybrid features.Finally,the affinity matching algorithm based on local features is employed to obtain the final retrieval results.Experiments demonstrate that the proposed attention-based cross-modal person reidentification algorithm was evaluated on the LLCM,SYSU-MM01,and RegDB datasets.It achieved Rank-1 accuracies of 71.0%,77.5%,and 95.0%,respectively.
【Key words】 Cross-modal person re-identification; Attention mechanisms; Multi-scale feature fusion; Feature alignment;
- 【网络出版投稿人】 北京邮电大学 【网络出版年期】2026年 01期
- 【分类号】TP391.41