节点文献

基于视频的小熊猫个体识别方法研究

Video-Based Individual Red Panda Recognition

【作者】 李蕾;

【导师】 赵启军; 黄珂;

【作者基本信息】 四川大学 , 电子信息(专业学位), 2024, 硕士

【摘要】 在当前全球环境变化和生物多样性丧失日益严峻的背景下,濒危物种的保护变得尤为重要。小熊猫(Ailurus fulgens)作为中国重要的珍稀濒危物种,其保护工作不仅关乎生物多样性的维护,也是生态文明建设的重要组成部分。近年来,随着人类活动的扩展和气候变化的影响,小熊猫面临着栖息地破坏、猎杀与贸易等多重威胁,种群数量持续下降,亟需有效的保护与研究手段。本文围绕小熊猫的个体识别技术展开研究,旨在通过最新的计算机视觉技术,提升对小熊猫种群监测和保护管理的效率和精度。传统的动物个体识别方法虽然在一定程度上实现了对动物个体的识别,但存在着人力成本高、识别准确率受限、无法应对复杂场景等问题。随着计算机视觉技术的快速发展,尤其是深度学习技术的应用,为动物个体识别提供了新的解决方案。但目前的研究仍面临一系列挑战:(1)基于特定部位的识别方法虽展现了其有效性,但在实际应用中,特定部位的清晰可见性难以保证,且对动物姿势变化和部位遮挡特别敏感。(2)基于全身图像的方法虽然能提供整体信息,但可能缺乏足够的细节信息,尤其是在复杂背景中。此外,无法使用单张图像来捕捉动态特征,如动物的行为模式,而这在区分相似个体时非常关键。(3)基于视频的方法虽然能够利用动态信息,但目前的动物个体识别方法并未利用到这些信息,且在实际应用中,如何有效地提取和利用时序信息仍然是一个挑战。本文针对以上挑战开展研究,主要工作如下:(1)构建了一个小熊猫视频数据库,其中包含60只小熊猫的1595段视频序列。在此基础上,引入了基于视频的行人重识别中的循环卷积神经网络、3D卷积神经网络、时间池化、时间注意力这四种时间建模方法,通过与图片基线方法的对比,证明了时序特征在小熊猫个体识别中的重要性。(2)利用时序信息进行运动特征提取。提出了一种基于运动特征聚合的视频小熊猫个体识别方法。在此方法中,通过一个运动注意力模块来捕获视频序列内的瞬时运动信息。该模块通过计算运动注意力权重来增强对运动特征的表示,同时利用残差连接保留外观信息。此外,该方法还融合了一个非局部特征融合模块,通过在ResNet-50结构内部集成非局部注意力机制,使模型能够捕获长期运动特征。在小熊猫视频数据集上的实验结果证明了该方法的有效性。(3)利用时序信息进行局部外观特征增强。针对小熊猫外观易出现的复杂形变和遮挡问题设计了一种基于局部外观特征增强的视频小熊猫个体识别方法。该方法包含判别性局部区域搜寻模块和局部区域超图模块,通过自适应学习判别性区域和建立时间依赖关系,进一步增强了小熊猫个体的特征表示能力。在小熊猫视频数据集上的实验结果证明了该方法是有效的。

【Abstract】 The protection of endangered species has become particularly important in the current context of global environmental change and the increasing loss of biodiversity.Among these species,the conservation of the red panda(Ailurus fulgens),a rare and endangered species en-demic to China,is not only crucial for maintaining biodiversity but also an important aspect of building ecological civilization.In recent years,the expansion of human activities and the impact of climate change have subjected the red panda to multiple threats,such as habitat de-struction,hunting,and trade,leading to a continuous decline in their population.This urgently necessitates effective conservation and research tools.This thesis focuses on the individual identification technology for the red panda,aiming to enhance the efficiency and accuracy of monitoring and conservation management of red panda populations through the latest advance-ments in computer vision technology.Although traditional methods of animal individual recognition achieve recognition to a certain extent,they suffer from problems such as high labor costs,limited recognition accu-racy,and an inability to cope with complex scenes.With the rapid development of computer vision technology,especially the application of deep learning,a new solution for animal indi-vidual recognition has emerged.However,current research still faces several challenges:(1)Recognition methods based on specific body parts have shown effectiveness,but the clear vis-ibility of these parts cannot always be guaranteed in practical applications.These methods are particularly sensitive to changes in the animal’s posture and to part occlusion.(2)Methods based on whole-body images can provide overall information but may lack detailed informa-tion,especially against complex backgrounds.Moreover,static images cannot capture dynamic features,such as behavioral patterns,which are critical for distinguishing similar individuals.(3)Although video-based methods can utilize dynamic information,current animal individual recognition methods do not effectively exploit this information,making it a challenge to extract and use temporal information in practical applications.This thesis addresses the aforementioned challenges,and the main contributions are as follows:(1)This thesis establishes a red panda video database,which encompasses 1,595 video se-quences of 60 individual red pandas.Building upon this foundation,it introduces four temporal modeling methods commonly employed in video-based person re-identification—namely,Re-current Neural Network(RNN),3D Convolutional Neural Network(3D CNN),Temporal Pool-ing,and Temporal Attention.By comparing these methods with the image baseline approach,it demonstrates the significance of temporal features in the individual recognition of red pandas.(2)Using temporal information for motion feature extraction,this thesis proposes an in-dividual recognition method for video red pandas based on motion feature aggregation.In this method,a motion attention module accurately captures instantaneous motion information within a video sequence.This module enhances the representation of motion features by computing motion attention weights,while retaining appearance information using residual concatenation.Additionally,the method incorporates a non-local feature fusion module that enables the model to capture long-term motion features by integrating a non-local attention mechanism within the ResNet-50 structure.Experimental results on the red panda video dataset demonstrate the effectiveness of this method.(3)Using temporal information for appearance feature enhancement,this thesis also de-signs a video red panda individual recognition method based on local appearance feature en-hancement to address the complex deformation and occlusion problems that red panda appear-ance is prone to.The method contains a discriminative local region search module and a hy-pergraph feature aggregation module,which further enhance the feature representation of red panda individuals by adaptively learning discriminative regions and establishing temporal de-pendencies.Experimental results on the red panda video dataset demonstrate the effectiveness of this method.

  • 【网络出版投稿人】 四川大学
  • 【网络出版年期】2025年 08期
  • 【分类号】TP391.41;Q958
节点文献中: 

本文链接的文献网络图示:

本文的引文网络