节点文献
ViT融合局部特征的无监督行人再识别研究
Local Feature Fusion based on Vision Transformer for Unsupervised Person Reidentification
【摘要】 目前无监督行人再识别是直接利用全局特征计算相似性,这会导致伪标签生成质量不佳,并且基于卷积神经网络的方法可能会产生细节丢失。针对这些问题,本文设计了一种基于局部特征融合的无监督行人再识别网络,通过随机滑窗的图像编码方式来解决信息丢失问题。为了产生可靠的伪标签,通过Vision Transformer独有的“抛弃特征”作为局部变量来计算局部相似性,并且融合摄像头相似性以提高总体样本相似性计算的准确性,提升伪标签生成的质量。实验结果表明,该方法在公开数据集Market-1501和DukeMTMC-reID上可以大幅提升模型性能。
【Abstract】 At present, unsupervised pedestrian recognition directly utilizes global features to calculate similarity,which can lead to poor quality of pseudo label generation, and methods based on convolutional neural networks may result in detail loss. In response to these issues, this article designs an unsupervised pedestrian recognition network based on local feature fusion, which solves the problem of information loss through random sliding window image encoding. In order to generate reliable pseudo labels, the unique "discarded features" of Vision Transformer are used as local variables to calculate local similarity, and camera similarity is fused to improve the accuracy of overall sample similarity calculation and improve the quality of pseudo label generation. The experimental results show that this method can significantly improve model performance on publicly available datasets Market-1501and DukeMTMC reID.
【Key words】 Person Re-identification; Vision Transformer; Unsupervised Learning; Clustering Algorithm;
- 【文献出处】 福建电脑 ,Journal of Fujian Computer , 编辑部邮箱 ,2023年12期
- 【分类号】TP391.41
- 【下载频次】115