节点文献

聚焦注意力与紧致特征融合Transformer的城市遥感影像语义分割

A Focused Attention and Feature Compact Fusion Transformer for Semantic Segmentation of Urban Remote Sensing Images

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 周国宇; 张菁; 闫伊; 卓力;

【Author】 ZHOU Guoyu;ZHANG Jing;YAN Yi;ZHUO Li;School of Information Science and Technology, Beijing University of Technology;Beijing Key Laboratory of Computational Intelligence and Intelligent System,Beijing University of Technology;

【通讯作者】 张菁;

【机构】 北京工业大学信息科学技术学院; 北京工业大学计算智能与智能系统北京市重点实验室;

【摘要】 在空天信息智能处理深度融合遥感数据获取与智能解译技术的推动下,城市遥感影像(URSI)语义分割逐渐发展成为连接空天信息与城市计算的关键研究方向。然而,与通用遥感影像相比,URSI中地物目标具有高度多样性和复杂性,表现为同类地物内部细节有差异、不同类地物之间特征相似易混淆,同时地物边界往往模糊且形态不规则,这些因素共同构成了其精细化分割面临的挑战。尽管基于Transformer的遥感影像语义分割方法取得了显著进展,但将其用于URSI时,不仅要考虑其对细节和边缘的提取能力,还需应对自注意力机制带来的计算复杂度等问题。为此,该文在编码器端引入聚焦注意力,以高效捕捉类内和类间关键特征;同时在解码器端对边缘特征进行紧致融合。针对URSI的独特特性,该文提出一种聚焦注意力与紧致特征融合Transformer语义分割模型(F3Former)。首先,在编码器端引入特征聚焦编码块(FFEB),通过建模Query-Key特征对的方向性,在保持较低线性复杂度的同时提升类内特征聚合与类间判别能力;在解码器端设计紧致特征融合模块(CFFM),结合深度卷积降低跨通道冗余计算,增强URSI边缘区域的细粒度分割表现。实验结果表明,该文提出的F3Former在Potsdam, Vaihingen和LoveDA数据集上的mIoU分别为88.33%, 81.32%和53.16%,计算成本减少到35.42 M Params, 48.02 GFLOPs和0.09 s测试时间,相较基线计算成本下降了28.91 M Params和194.86 GFLOPs,显著平衡了URSI语义分割的精度和速度。

【Abstract】 Objective Driven by the growing integration of remote sensing data acquisition and intelligent interpretation technologies within aerospace information intelligent processing, semantic segmentation of Urban Remote Sensing Image(URSI) has emerged as a key research area connecting aerospace information and urban computing. However, compared to general Remote Sensing Image(RSI), URSI exhibits a high diversity and complex of geo-objects, characterized by fine-grained intra-class variations, inter-class similarities that cause confusion, as well as blurred and irregular object boundaries. These factors present difficulties for fine-grained segmentation. Despite their success in RSI semantic segmentation, applying Transformer-based methods to URSI requires a balance between capturing detailed features and boundaries, and managing the computational cost of self-attention. To address these issues, this paper introduces a focused attention mechanism in the encoder to efficiently capture discriminative intra-and inter-class features, while performing compact edge feature fusion in the decoder.Methods This paper proposes a Focused attention and Feature compact Fusion Transformer(F3 Former). The encoder incorporates a dedicated Feature-Focused Encoding Block(FFEB). By leveraging the focused attention mechanism, it adjusts the directions of Query and Key features such that the features of the same class are pulled closer while those of different classes are repelled, thereby enhancing intra-class consistency and interclass separability during feature representation. This process yields a compact and highly discriminative attention distribution, which amplifies semantically critical features while curbing computational overhead. To complement this design, the decoder employs a Compact Feature Fusion Module(CFFM), where Depth-Wise Convolution(DW Conv) is utilized to minimize redundant cross-channel computations. This design strengths the discriminative power of edge representations, improvs inference efficiency and deployment adaptability, and maintains segmentation accuracy.Results and Discussions F3 Former demonstrates favorable performance on several benchmark datasets,alongside a lower computational complexity. On the Potsdam and Vaihingen benchmarks, it attained mIoU scores of 88.33%/81.32%, respectively, ranking second only to TEFormer with marginal differences in accuracy(Table 1). Compared to other lightweight models including CMTFNet, ESST, and FSegNet, F3 Former consistently delivered superior results in mIoU, mF1, and PA, demonstrating the efficacy of the proposed FFEB and CFFM modules in capturing complex URSI features. On the LoveDA dataset, it reached 53.16% mIoU and outperformed D2 SFormer in several critical categories(Fig. 4). Moreover, F3 Former strikes a favorable balance between accuracy and efficiency, reducing parameter count and FLOPs by over 30% compared to TEFormer,with only negligible degradation in accuracy(Table 2). Qualitative results further indicate clearer boundary delineation and improved recognition of small or occluded objects relative to other lightweight approaches(Fig. 5 and Fig. 6). Ablation studies validate the critical role of both the Focused Attention(FA) mechanism and the Compact Feature Fusion Head(CFFHead) in achieving accuracy and efficiency gains(Tables 3 & Tables 4).Conclusions This work tackles key challenges in URSI semantic segmentation—including intra-class variability, inter-class ambiguity, and complex boundaries—by proposing F3 Former. In the encoder, the FFEB improves intra-class aggregation and inter-class discrimination through directional feature modeling. In the decoder, the CFFM employs DW Conv to minimize redundancy and enhance boundary representations. With linear complexity, F3 Former attains higher accuracy and stronger representational capacity while remaining efficient and deployment-friendly. Extensive experiments across multiple URSI benchmarks confirm its superior performance, highlighting its practicality for large-scale URSI applications. However, compared to existing State-Of-The-Art(SOTA) lightweight methods, the computational efficiency of the FFEB still has room for improvement. Future work is directed towards replacing Softmax with a more efficient operator to accelerate attention computation, maintaining accuracy while advancing efficient URSI semantic segmentation.Additionally, as the decoder’s channel interaction mechanism remains relatively limited, we plan to incorporate lightweight attention or pointwise convolution designs to further strengthen feature fusion.

【基金】 北京市自然科学基金(L247025)~~
  • 【文献出处】 电子与信息学报 ,Journal of Electronics & Information Technology , 编辑部邮箱 ,2025年12期
  • 【分类号】TP751
  • 【下载频次】26
节点文献中: 

本文链接的文献网络图示:

本文的引文网络