节点文献

面向无人机航摄图像语义分割的双路特征融合网络

Dual-Stream Feature Aggregation Network for Unmanned Aerial Vehicle Aerial Images Semantic Segmentation

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李润增; 史再峰; 孔凡宁; 赵向阳; 罗韬;

【Author】 Li Runzeng;Shi Zaifeng;Kong Fanning;Zhao Xiangyang;Luo Tao;School of Microelectronics, Tianjin University;Tianjin Key Laboratory of Imaging and Sensing Microelectronic Technology;College of Intelligence and Computing, Tianjin University;

【通讯作者】 史再峰;

【机构】 天津大学微电子学院; 天津市成像与感知微电子技术重点实验室; 天津大学智能与计算学部;

【摘要】 针对无人机航摄图像中目标尺寸差异大导致的感受野难以同时兼顾不同尺寸物体分割效果的问题,提出了利用两路分支分别提取浅层和深层信息的双路特征融合网络(DSFA-Net)。在编码器中,浅层分支利用三个串行ConvNeXt模块提取高通道数的浅层特征以保留更多空间细节;深层分支利用坐标注意力空洞空间金字塔池化(CA-ASPP)模块为特征图重新分配权重,使网络更加关注尺寸各异的分割目标,获得深层多尺度特征。在解码过程中,网络利用双边引导融合模块为两层特征建立通信以进行分辨率融合,提高层级特征的利用率。所提方法在AeroScapes和Semantic Drone航摄图像数据集上进行了实验,其平均交并比分别达到83.16%和72.09%、平均像素准确率分别达到90.75%和80.34%。与主流的语义分割方法相比,所提方法对于具有较大尺寸差异的目标,分割能力更强,更适用于无人机航摄图像场景下的语义分割任务。

【Abstract】 Large object size difference in unmanned aerial vehicle(UAV) aerial photography makes it difficult to take into account the segmentation effect of objects of different sizes in the receptive field. A dual-stream feature aggregation network(DSFA-Net)with two branches to extract low-level and high-level features separately, is proposed for such problems. In the encoder, a lowlevel information extraction branch with three serial ConvNeXt modules is used to preserve more low-level features by generating more channels of features. In the deep feature branch, the coordinate attention atrous spatial pyramid pooling(CA-ASPP) module reassigns weights to feature maps in the channel dimension. It makes the module focus on segmentation objects of different sizes and deep-level multi-scale features are obtained. During the decoding process, the bilateral guided aggregation module performs resolution aggregation between the low-level and deep-level features. Our method is evaluated on the AeroScapes and Semantic Drone datasets, the mean intersection over union is 83. 16% and 72. 09% respectively, and the mean pixel accuracy is 90. 75%and 80. 34% respectively. The proposed method is more capable of segmenting objects with large difference sizes compared to mainstream methods. It is suitable for semantic segmentation tasks for UAV aerial images.

【基金】 国家自然科学基金(62071326);天津市自然科学基金(22JCYBJC00140)
  • 【文献出处】 激光与光电子学进展 ,Laser & Optoelectronics Progress , 编辑部邮箱 ,2023年24期
  • 【分类号】P231;TP391.41
  • 【下载频次】152
节点文献中: 

本文链接的文献网络图示:

本文的引文网络