节点文献

监控视频中吸烟行为检测方法研究

Research on Smoking Behavior Detection Algorithm in Surveillance Video

【作者】 张志伟;

【导师】 路小波;

【作者基本信息】 东南大学 , 控制科学与工程, 2025, 硕士

【摘要】 吸烟行为不但对人体健康造成直接危害,还会造成环境污染、火灾隐患等问题。在监控视频情景中,对吸烟行为进行准确、及时的检测对于当前无烟环境建设工作的推进具有重要意义。由于监控视频中作为吸烟行为判据的香烟目标具有目标小、易受干扰的特点,本文对监控视频情景中的吸烟检测方法进行了技术研究,主要包括以下四个方面的内容:一、基于多尺度注意力机制的香烟目标检测模型。针对YOLOv8的特征流向不利于特征提取的问题,在Neck部分设计了特征信息反转的FPPAN网络,优化了小目标特征提取能力;针对现有的信息提取模块在多尺度信息提取和融合方面能力较弱的问题,设计了MSP瓶颈机制,利用多尺度的特征融合提高了多尺度信息的处理能力;针对现有大量级联模块会造成过多参数量提升的问题,设计了MSP特征提取单元,通过特征分离并处理的方式,避免了全局信息处理造成的额外计算量提升。在实验中,和其它现有方法对比验证了改进能够达成精度和效率的平衡。二、基于多模态困难负样本挖掘的误检抑制方法。针对现有目标检测神经网络误检率较高的问题,设计了负样本挖掘迭代优化方法,在迭代训练中提取检测负样本,抑制目标检测神经网络在背景区域的误检测现象。针对现有负样本挖掘方法的语义没有与检测网络对齐的问题,提出了基于多模态和尺度的错误样本分析网络,聚焦关键判据;针交叉熵损失难以保证样本特征区分度的问题,提出了基于改进对比度损失的损失函数设计,取得语义提取能力和样本区分度之间的平衡。最终设计了实验验证了方法的有效性。三、融合多模态与运动约束建模的增强检测方法。针对单帧视频检测模型无法融合时空特征的问题,提出了一套融合多模态特征增强与运动约束建模的实时检测方法。引入ROI Align机制、构建异构通道对齐压缩模块,以及提出双流上下文注意力增强模块,提升了特征表达对形变、模糊及局部遮挡的鲁棒性;提出基于长轴角速度约束的改进模型,显式建模香烟目标的旋转运动特性,并设计动态置信度衰减策略,在保证推理速度的同时维护过往目标信息。通过实验验证了设计的有效性。四、吸烟行为检测软件的设计与实现。使用Python语言,对吸烟行为检测软件进行了设计与实现。设计并实现了通用处理模块、香烟检测模块以及跟踪和重识别模块。在指定的部署平台上,通过软件实验测试验证其能够满足实际应用的需要。

【Abstract】 Smoking not only directly harms human health,but also causes environmental pollution,fire hazards,and other problems.In surveillance video scenarios,accurate and timely detection of smoking behavior is of great significance for promoting the construction of smoke-free environments.Since cigarette targets,which serve as criteria for judging smoking behavior,are small and susceptible to interference in surveillance videos,this paper conducts technical research on smoking detection methods in surveillance video scenarios,mainly including the following four aspects:1.A cigarette target detection model based on a multi-scale attention mechanism.To address the problem that the feature flow of the existing YOLOv8 is not conducive to feature extraction,an FPPAN network with reversed feature information is designed in the Neck part to optimize the feature extraction ability of small targets.To address the problem that existing information extraction modules have weak ability in multi-scale information extraction and fusion,an MSP bottleneck mechanism is designed,which improves the processing ability of multi-scale information by utilizing multi-scale feature fusion.To address the problem that existing large-scale cascade modules cause excessive parameter increases,an MSP feature extraction unit is designed,which avoids the additional computational burden caused by global information processing by separating and processing features.In experiments,comparisons with other existing methods verify that the improvements can achieve a balance between accuracy and efficiency.2.A false detection suppression framework based on cross-modal hard negative sample mining.To address the problem of high false detection rate in existing object detection neural networks,a negative sample mining iterative optimization framework is designed,which extracts detection negative samples in iterative training and suppresses false detection phenomena of object detection neural networks in background regions.To address the problem that the semantics of existing negative sample mining methods are not aligned with the detection network,a cross-modal and scale-based error sample analysis network is proposed,focusing on key criteria.To address the problem that cross-entropy loss is difficult to guarantee the feature distinguishability of samples,a loss function design based on improved contrastive loss is proposed,achieving a balance between semantic extraction ability and sample distinguishability.Finally,experiments are designed to verify the effectiveness of the method.3.An enhanced re-detection framework integrating multi-scale perception and motion constraint modeling.To address the problem that single-frame video detection models cannot integrate spatiotemporal features,a real-time tracking framework integrating multi-modal feature enhancement and motion constraint modeling is proposed.ROI Align mechanism is introduced,a heterogeneous channel alignment compression module is constructed,and a dual-stream context attention enhancement module is proposed to improve the robustness of feature representation to deformation,blur,and partial occlusion.An improved model based on long-axis angular velocity constraint is proposed to explicitly model the rotational motion characteristics of cigarette targets,and a dynamic confidence decay strategy is designed to maintain past target information while ensuring inference speed.Experiments demonstrate the effectiveness of the designed framework.4.Design and implementation of smoking behavior detection software.Using Python language,the smoking behavior detection software was designed and implemented.A general processing module,a cigarette detection module,and a tracking and re-identification module were designed and implemented.On the specified deployment platform,software experiments and tests verified that it can meet the needs of practical applications.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2026年 07期
  • 【分类号】TP391.41;TP311.52
节点文献中: