节点文献
密集场景下目标计数与定位研究
Object Counting and Localization Research in Dense Scenarios
【作者】 李波;
【导师】 张勇;
【作者基本信息】 北京工业大学 , 电子信息(专业学位), 2024, 硕士
【摘要】 随着社会和科技的不断进步,对图像和视频中目标的精确计数与定位的需求日益增长。这一需求在智能交通、医学影像分析、生态环境监测等多个领域尤为显著,准确掌握目标数量和位置对于深入解析复杂场景、增强安全防护以及提升社会运行效率具有至关重要的作用。例如,在智能交通系统中,通过精确地计数和定位行人和车辆,可以显著提高交通管理的效率,并有效减少交通事故。因此,开发高效且准确的目标计数与定位技术,对于解决实际问题、提高工作效率以及推动社会发展均具有重大意义。本研究通过结合先进的图像处理技术和深度学习算法,旨在实现图像和视频中目标的精确计数与定位,涵盖以下几个主要研究方向:首先,目标计数任务中的不均匀密度分布是导致计数模型性能下降的一个关键因素。在强监督目标计数方法中,利用目标位置信息生成的密度图已成功地引入了目标分布信息,显著提升了计数性能。然而,在弱监督计数场景中,缺乏准确的位置信息使得模型难以精准捕捉目标的密度分布,这极大地限制了模型的效果。为应对这一挑战,本研究提出了一种创新的超图关联模块。该模块利用超图神经网络来建模特征之间的复杂关系,从而提升模型对场景的理解能力。通过这种方式,模型能够获取更为丰富和准确的目标密度分布信息,从而显著提升计数精度。特别是在人群计数场景中,考虑到人群聚集特性和密集场景中目标密度分布的极端不均匀性,本文特别针对这些情况进行了实验验证,展示了所提方法的有效性和优越性。其次,现有目标定位方法主要通过将复杂多样的目标准确映射到相应的定位图上,然后处理这些定位图以获取精确的位置信息。然而,由于目标在形状、尺度和颜色上的剧烈变化,实现一致性的映射面临着巨大的挑战。这一问题在医学图像的细胞定位任务中尤为明显,因此该任务成为了本研究的主要验证场景。为了克服这一难题,本研究首次将定位任务转变为图像与定位图之间的特征对齐任务,并提出了一个多尺度超图模块,该模块能够统一处理定位任务中由形状、尺度和颜色剧烈变化带来的挑战。通过这种创新方法,显著提高了目标定位的精度,并为处理目标多样性问题提供了新的思路和技术手段。最后,本研究首次提出了小样本目标定位任务,并开发了一个高性能的小样本目标定位框架。该任务通过对查询图像中目标类别的支持性描述,实现了对查询图像中目标的精确定位。面对查询图像中目标类内差异大和定位过程中可能出现的目标漏检两大挑战,本研究设计了双通路特征增强模块。该模块有效地建立了支持图像和查询图像间的特征关联,提高了目标间的区分度,从而大幅提升了定位性能。此外,研究中还引入了自查询模块,该模块利用查询图像中的分布信息来优化最终定位图的生成,为小样本目标定位任务的研究和实践设定了高性能基准。通过这一框架,模型能够有效应对未见过的目标类别,显著拓宽了目标定位应用的领域。通过上述创新性方法,本研究为密集场景下的目标计数与定位任务提供了更加准确和鲁棒的解决方案。这些解决方案不仅适用于特定场景,还能够被广泛应用于更多的任务场景中。特别是在城市管理和社会安全领域,这些技术的应用将极大地促进监控、交通管理和环境保护等方面的工作,从而推动城市的智能化发展和社会的整体进步。
【Abstract】 As society and technology continue to advance,the demand for precise counting and localization of objects in images and videos is growing.This demand is particularly notable in fields such as intelligent transportation,medical image analysis,and ecological environment monitoring.Precise localization and counting are crucial for analyzing complex scenes in depth,enhancing safety measures,and improving the efficiency of societal operations.For example,in intelligent transportation systems,by precisely counting and locating pedestrians and vehicles,traffic management efficiency can be significantly enhanced,and traffic accidents can be effectively reduced.Therefore,developing efficient and accurate techniques for object counting and localization is of great importance for solving practical problems,enhancing work efficiency,and driving social development.This study aims to achieve precise counting and localization of objects in images and videos by integrating advanced image processing technologies and deep learning algorithms,covering the following main research directions:First,the uneven density distribution in object counting tasks is a key factor leading to the decline in model performance.In strongly supervised object counting methods,density maps generated using object location information have successfully incorporated objects distribution information,significantly enhancing counting performance.However,in weakly supervised counting scenarios,the lack of accurate location information makes it difficult for models to accurately capture the density distribution of objects,greatly limiting the effectiveness of the models.To address this challenge,this study proposes an innovative hypergraph association module.This module utilizes a hypergraph neural network to model complex relationships among features,thereby enhancing the model’s ability to understand scenes.In this way,the model can obtain richer and more accurate object density distribution information,significantly improving counting accuracy.Especially in crowd counting scenarios,considering the aggregation characteristics of crowds and the extreme unevenness of object density distribution in dense scenes,this paper specifically conducts experimental verification for these situations,demonstrating the effectiveness and superiority of the proposed method.Secondly,existing object localization methods mainly involve accurately mapping complex and diverse objects onto corresponding location maps,and then processing these maps to obtain precise location information.However,due to drastic variations in shape,scale,and color of the objects,achieving consistent mapping faces significant challenges.This issue is particularly evident in the cell localization tasks of medical imaging,thus making it a primary validation scenario in this study.To overcome this challenge,this research for the first time transforms the localization task into a feature alignment task between images and localization maps,and introduces a multi-scale hypergraph module.This module can uniformly address the challenges brought about by severe changes in shape,scale,and color in localization tasks.Through this innovative approach,the accuracy of object localization has been significantly improved,providing new ideas and technical means for dealing with the diversity of objects.Finally,this study pioneers the task of few-shot object localization and develops a high-performance framework for it.The task achieves precise localization of objects in query images by using supportive descriptions of object categories.Faced with the challenges of significant intra-class variance and potential objects omissions during the localization process,the study introduces a dual-path feature enhancement module.This module effectively establishes feature associations between support and query images,enhancing objects differentiation,and significantly improving localization performance.Additionally,a self query module is introduced,leveraging distribution information from query images to optimize the creation of the final location map,setting a high-performance benchmark for the few-shot object localization task.This framework enables the model to effectively handle unseen object categories,substantially broadening the application scope of object localization.Through these innovative methods,this study provides more accurate and robust solutions for object counting and localization tasks in dense scenarios.These solutions are not only suitable for specific scenarios but can also be broadly applied to a wider range of task scenarios.Particularly in the fields of urban management and public safety,the application of these technologies will greatly enhance efforts in monitoring,traffic management,and environmental protection,thereby promoting the intelligent development of cities and overall societal progress.
【Key words】 object counting; object localization; hypergraph neural networks; few-shot object localization;
- 【网络出版投稿人】 北京工业大学 【网络出版年期】2025年 07期
- 【分类号】TP391.41