节点文献
融合可形变卷积与注意力检测头的交通多目标检测
Traffic multi-object detection integrating deformable convolution and attention detection head
【摘要】 目标检测作为环境感知的关键环节,在复杂交通场景中需应对多尺度目标识别和目标严重遮挡等挑战,这些问题往往会导致YOLO模型边界框定位精度下降,最终出现漏检和误检的情况。为应对上述挑战,设计并改进了基于YOLOv11架构的多目标检测方案:设计矩形自校准扩张多尺度融合模块(DMSFM),通过空洞卷积与多尺度融合策略实现三重特征提取与融合,提升模型对不同尺寸目标的特征感知能力;引入Shape-NWD损失函数,突破传统IoU的局限,以形状加权和归一化Wasserstein距离的几何匹配准则,优化不同尺寸目标锚框的定位精度与敏感度;融合含并行补丁感知注意力机制的检测头,通过多分支与注意力策略,强化模型对多尺度目标特征的适应性及分类决策能力。在CODA自动驾驶道路目标检测数据集上的实验结果表明,所提方法相较于基准模型,平均召回率相对提升26.5%,mAP50和mAP50-95分别相对提升24.1%和16.9%;消融实验验证了三个核心组件的有效协同,进一步证实了本方案在复杂交通场景多尺度目标检测任务中的高鲁棒性。
【Abstract】 Object detection, as a critical component of environmental perception, faces challenges in complex traffic scenarios such as multi-scale object recognition and severe occlusion. These issues often lead to reduced bounding box localization accuracy in YOLO models, ultimately resulting in missed detections and false positives. To address this challenge, this paper designs and improves a multi-object detection solution based on the YOLOv11 architecture. Designing a rectangular self-calibrating expansion multi-scale fusion module(DMSFM), which achieves triple feature extraction fusion through dilated convolutions and multi-scale fusion strategies, thereby enhancing the model feature perception for targets of different sizes. The Shape-NWD loss function is introduced to overcome the limitations of traditional IoU, it employs a geometric matching criterion based on shape-weighted and normalized Wasserstein distance to optimize the localization accuracy and sensitivity of anchor boxes f for targets of different sizes. The detection head integrates a parallel patch-aware attention mechanism, enhancing the model adaptability to multi-scale target features and classification decision-making capabilities through multi-branch and attention strategies. The experimental results on the CODA autonomous driving road object detection dataset demonstrate that the proposed method achieves an average recall rate improvement of 26.5% compared to the benchmark model, with relative improvements of 24.1% and 16.9% in mAP50 and mAP50-95, respectively. The ablation experiment verifies the effective collaboration of the three core components, further confirming the high robustness of the proposed approach in multi-scale object detection tasks for complex traffic scenarios.
【Key words】 object detection; multi-scale integration; dilated convolution; Shape-NWD; DMSFM;
- 【文献出处】 江苏理工学院学报 ,Journal of Jiangsu University of Technology , 编辑部邮箱 ,2026年01期
- 【分类号】TP391.41;U463.6
- 【下载频次】49