节点文献

高分辨率遥感影像建筑物提取方法:轻量级全局注意力网络模型U~2-former

A Lightweight U~2-Former with Global Attention for Building Extraction from High Resolution Remote Sensing Images

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘文璐蔡玉林禚越常志鹏赵相伟

【Author】 LIU Wenlu;CAI Yulin;ZHUO Yue;CHANG Zhipeng;ZHAO Xiangwei;College of Geomatics and Spatial Information,Shandong University of Science and Technology;

【通讯作者】 蔡玉林;

【机构】 山东科技大学测绘与空间信息学院

【摘要】 【目的】建筑物信息对于城市规划、环境监测、灾害应急等领域具有重要应用价值。随着高分辨率遥感影像的普及,如何高效、精确地实现建筑物自动提取已成为遥感领域的研究热点。尽管深度学习方法提升了遥感影像建筑物提取的效率与精度,但现有算法在轮廓完整性、抗干扰能力和模型轻量化方面仍面临挑战。本文旨在通过融合卷积与自注意力机制的优势,构建一种兼顾高精度与低复杂度的建筑物提取网络,以推动该方法在资源受限环境下的实际应用。【方法】为解决这些问题,本文将U~2-Net网络与Transformer模块相融合,提出了一种新的轻量化全局注意力网络模型U~2-former。该方法在U~2-Net网络基础上采取了3个方面的改进:(1)在编码部分引入通道注意力机制来强化局部特征捕获能力;(2)将解码器重构为Transformer模块,通过优化的多头注意力机制与通道增强型多层感知机建立全局空间依赖;(3)采用多级特征融合策略整合不同解码层输出,以提升边界完整性。【结果】在WHU航空影像、Massachusetts及Inria航空图像三大基准数据集上的实验结果表明:U~2-former模型仅有6 M参数量,但IoU指数仍然分别达到了91.69%、74.96%和80.13%的精度,优于当前的主流深度学习算法;消融实验进一步证明了3种改进措施的有效性:相对于U~2-Net,所提出的U~2-former模型在3种数据集上的IoU值相比U~2-Net分别提升了1.33%、3.54%和2.24%。同时,可视化结果进一步显示,U~2-former在复杂场景中能有效保持建筑物轮廓完整性,减少错检、漏检,尤其对阴影遮挡、结构复杂建筑具有较强的鲁棒性。【结论】U~2-former通过融合CNN的局部特征提取能力与Transformer的全局上下文建模优势,在较低参数量下实现了高精度的建筑物提取。该方法有效解决了轻量化与特征完整性、局部细节与全局依赖、边界精度与抗干扰能力之间的关键矛盾,不仅在高分辨率遥感影像建筑物提取任务中表现优异,也为边缘计算等资源受限场景下的遥感影像智能解译提供了新的技术路径。

【Abstract】 [Objectives] Building information plays a crucial role in applications such as urban planning, environmental monitoring, and disaster emergency response. With the widespread adoption of high-resolution remote sensing imagery, how to achieve efficient and accurate automatic building extraction has become a key research focus in the field of remote sensing. While deep learning approaches have enhanced the efficiency and precision of building extraction from remote sensing images, existing algorithms continue to encounter challenges pertaining to contour integrity, anti-interference capability, and model lightweighting. The aim of this paper is to construct a building extraction network that balances high accuracy and low complexity by fusing the advantages of convolution and self-attention mechanisms, in order to promote the practical application of this method in resource-constrained environments. [Methods] To tackle these issues, this study integrates the U~2-Net network with Transformer modules, thereby proposing a novel lightweight global attention network model termed U~2-former. This method introduces three improvements based on the U~2-Net network:(1) the incorporation of a channel attention mechanism in the encoding segment to strengthen the capacity for capturing local features;(2) the reconstruction of the decoder into a Transformer module, with global spatial dependencies established via an optimized multihead attention mechanism and a channel-enhanced multi-layer perceptron;(3) the adoption of a multi-level feature fusion strategy to integrate outputs from different decoding layers, so as to improve boundary integrity. [Results] Experiments on three benchmark datasets, namely WHU aerial images, Massachusetts aerial images, and Inria aerial images, show that the U~2-former model has only 6 M parameters, yet its IoU indices still reach 91.69%, 74.96%, and 80.13% respectively, which outperforms those of current popular algorithms; The ablation experiments further demonstrate the effectiveness of the three improvements: relative to U~2-Net, the proposed U~2-former model improves the Io U values on the three datasets by 1.33%, 3.54%, and 2.24%, respectively. In addition, visualization results confirm that U~2-former effectively maintains building contour integrity in complex scenarios, reduces false detections and omissions, and exhibits strong robustness against shadow occlusion and structurally complex buildings. [Conclusions] By combining the local feature extraction capability of CNN with the global contextual modeling ability of Transformer, the U~2-former model achieves high-precision building extraction with very low parameter complexity. The proposed method effectively addresses the key trade-offs among model lightness, feature integrity, local detail preservation, global dependency modeling, and anti-interference capability. It not only achieves outstanding performance in building extraction from high-resolution remote sensing images but also offers a promising technical pathway for intelligent interpretation of remote sensing imagery in resourceconstrained scenarios such as edge computing.

【基金】 国家自然科学基金项目(42501483);山东省自然科学基金项目(ZR2022MD018);青岛市自然科学基金青年项目(25-1-1-50-zyyd-jch)~~
  • 【文献出处】 地球信息科学学报 ,Journal of Geo-information Science , 编辑部邮箱 ,2026年02期
  • 【分类号】TP18;TP751;P237
  • 【下载频次】27
节点文献中: 

本文链接的文献网络图示:

本文的引文网络