节点文献

基于特征净化与增强的跨模态行人再识别方法研究

Research on Cross-Modal Person Re-identification Method Based on Feature Purification and Enhancement

【作者】 杨涛

【导师】 梁莉莉;

【作者基本信息】 西安理工大学 , 控制科学与工程, 2025, 硕士

【摘要】 基于可见光和红外图像的跨模态行人再识别技术(Cross-modal Person Re-Identification,CM-ReID能够实现全天候的行人检索,在智慧家居、智能监控和智能交通等场景中具有广泛的应用前景。因此,对CM-ReID技术的研究具有重要的理论价值与实际意义。然而,现有方法在建模过程中存在对模态共享特征挖掘不充分、忽略模态特有特征的问题,从而限制了跨模态行人再识别模型性能的进一步提升。为此,本文对可见光与红外模态的特有特征净化与共享特征增强展开研究,具体研究内容如下:1)针对现有跨模态行人再识别方法容易忽略了模态特有特征中蕴含的关键判别信息,导致模型在缩小可见光与红外图像模态差距方面效果有限。本文设计了一种融合特征净化器的单双流网络结构,分别针对可见光图像与红外图像单独提取特征。利用模态鉴别器在保留可见光图像与红外图像模态差异的同时学习模态特有特征,同时借助模态混淆器削弱模态信息,引导生成可匹配的共享特征。此外,为抑制模态特有特征中冗余信息的干扰,实现对模态特有特征的净化,进一步引入融合实例规范化与混合域注意力机制的特征净化器模块。实例规范化用于对齐模态特征分布,混合域注意力机制则用于引导网络关注判别区域。最终输出更加清洁、可区分的跨模态特有特征,有效提升识别精度。为了证明该策略的有效性,在SYSU-MM01数据集上进行了验证实验:所提出方法在Rank-1识别准确率方面达到82.13%,mAP为75.75%,相比于无融合或无净化器的传统方法有明显提升,充分说明了所提出方法的有效性。2)针对现有方法中跨模态共享特征表达能力不足的问题,本文在前述网络基础上引入了基于TGA与CSA的特征增强模块。其中,TGA机制通过构建三重图结构,引导模型从空间与时序关系的角度进行语义结构对齐,通过结构对齐将模态特异性知识注入共享分支,有效缓解模态差异;CSA机制则引入类语义注意力模块,在语义层面执行类语义对齐操作,使模态特有与共享特征在同一类别语义空间内保持一致性。通过两者的协同作用,可以将净化后的模态特有特征中强判别性信息集成到模态共享特征中,并且模型在对行人行人语义关系建模、局部信息细节表达以及跨模态特征对齐等方面能力得到提升。同时构建由基础损失、实例规范化损失、TGA损失和CSA损失组成的综合损失函数,从多层面协同引导模型训练。实验结果表明,在SYSU-MM01数据集上,加入TGA与CSA机制后,模型的Rank-1识别率显著提高至84.98%,mAP也提升至77.29%,在多个评估指标上均优于现有主流方法,验证了本研究提出方法可以增强模型的判别性与鲁棒性。

【Abstract】 Cross-Modal Person Re-Identification(CM-ReID)based on visible and infrared images enables all-weather pedestrian retrieval,holding broad application prospects in scenarios such as smart homes,intelligent monitoring,and intelligent transportation.Therefore,the research on CM-ReID technology carries significant theoretical value and practical significance.However,existing methods suffer from insufficient mining of modality-shared features and neglect of modality-specific features during modeling,which limits the further improvement of cross-modal person re-identification model performance.To address this,this thesis investigates the purification of modality-specific features and enhancement of shared features between visible and infrared modalities,with the specific research contents as follows:1)Aiming at the problem that existing cross-modal person re-identification methods easily overlook the key discriminative information contained in modality-specific features,resulting in limited effectiveness in reducing the modality gap between visible and infrared images,this thesis designs a single-dual-stream network structure integrated with feature purifier to extract features from visible and infrared images separately.A modality discriminator is employed to learn modality-specific features while preserving the modality differences between visible and infrared images,while a modality mixer weakens modality information to guide the generation of matchable shared features.Additionally,to suppress the interference of redundant information in modality-specific features and achieve their purification,feature purifier module fusing instance normalization and hybrid-domain attention mechanism is further introduced.Instance normalization is used to align modality feature distributions,while the hybrid-domain attention mechanism guides the network to focus on discriminative regions.The final output is cleaner and more distinguishable cross-modal specific features,effectively improving recognition accuracy.Validation experiments on the SYSU-MM01 dataset demonstrate the effectiveness of this strategy:the proposed method achieves an 82.13%Rank-1 recognition significantly outperforming traditional methods without fusion or purifiers,fully illustrating its effectiveness.2)To address the problem of insufficient cross-modal shared features representation capability in existing methods,this thesis introduces a feature enhancement module based on TGA(Triple Graph Alignment)and CSA(Class Semantic Alignment)on the aforementioned network.The TGA mechanism constructs a triple graph structure to guide the model to perform semantic structure alignment from the perspectives of spatial and temporal relationships,injecting modality-specific knowledge into the shared branch through structure alignment to effectively alleviate modality differences.The CSA mechanism introduces a class semantic attention module to perform class semantic alignment at the semantic level,ensuring consistency between modality-specific and shared features within the same class semantic space.Through their synergistic effect,the highly discriminative information from the purified modality-specific features is integrated into the modality-shared features,and the model’s capabilities in modeling pedestrian semantic relationships,expressing local information details,and cross-modal feature alignment are enhanced.Meanwhile,a comprehensive loss function comprising basic loss,instance normalization loss,TGA loss,and CSA loss is constructed to synergistically guide model training from multiple levels.Experimental results on the SYSU-MM01 dataset show that after incorporating the TGA and CSA mechanisms,the model’s Rank-1 recognition rate significantly increases to 84.98%,and the mAP rises to 77.29%,outperforming existing mainstream methods in multiple evaluation metrics.This validates that the proposed method can enhance the model’s discriminability and robustness.

  • 【分类号】TP391.41
节点文献中: