节点文献
拉普拉斯增强与跨尺度显著性目标检测网络
Laplace-Enhanced Cross-Scale Salient Object Detection Network
【摘要】 针对现有显著性目标检测方法难以协调目标整体结构一致性与局部细节丰富性的问题,提出了基于拉普拉斯增强与跨尺度融合的显著性目标检测算法。首先,通过顺序双尺度编码器,将高分辨率图像及其下采样图输入权值共享的Swin Transformer骨干网络,以提取具有不同感受野的多层次特征,从而增强模型对全局语义的理解能力。其次,引入拉普拉斯算子生成高频特征图,并通过注意力机制将其与主干特征融合,提升模型对关键边缘和细节特征的感知能力。最后,设计了跨尺度特征聚合模块,通过串联残差的特征金字塔,实现跨尺度信息的高效融合。在五个标准数据集上的实验结果表明,所提方法显著优于当前主流方法,充分验证了所提模型在协调细节与结构建模方面的有效性及其优越性能。
【Abstract】 Objective Salient object detection(SOD), a pivotal pre-processing step for various computer vision tasks such as image segmentation and visual tracking, aims to identify the most visually distinctive objects in an image. While deep learning-based methods have achieved remarkable progress, a fundamental challenge persists: effectively balancing the preservation of local finegrained details(e.g., edges, textures) with the maintenance of global structural consistency of target objects. Many existing models, whether employing heterogeneous dual-branch networks, integrating edge supervision, or utilizing parallel multi-scale fusion, often struggle with this trade-off. They tend to either lose intricate details when capturing holistic semantics or produce structurally incoherent predictions when focusing on local features. This work is motivated by the necessity to bridge this gap. We propose a novel network architecture designed to harmonize detail richness and structural integrity from the feature extraction stage through to the final prediction, thereby advancing the robustness and accuracy of SOD in complex real-world scenarios.Methods We present the Laplace-enhanced scale-fusion network(LSFNet), which is built upon a classic encoder-neck-decoder framework with three key innovations. First, a progressive dual-scale encoder(PDE) is designed to replace common heterogeneous dual-branch encoders. It sequentially processes a half-resolution image and the original high-resolution image through a single weightshared Swin Transformer backbone. This strategy allows the low-resolution branch to first capture global semantic context, which then implicitly guides the subsequent high-resolution branch to focus on extracting complementary local details, ensuring semantic consistency and parameter efficiency. Second, an edge-guided joint modulation(EJM) module is introduced in the network neck to enhance boundary perception without relying on extra edge annotations. A fixed Laplacian convolution kernel is applied to the input image in an unsupervised manner to generate high-frequency edge features. These edge features are then used as spatial attention weights to modulate the fused multi-scale features from the PDE, selectively highlighting crucial boundary regions and suppressing feature dispersion. Third, a cross-scale feature aggregation(CFA) module is constructed within the decoder to overcome the limitations of simple parallel or additive fusion. It is based on a serial residual pyramid pooling(SRPP) structure. The module takes features from different layers and employs an implicit detail-and-global enhancement mechanism followed by a cascade of dilated convolutions with progressively increasing dilation rates(e. g., [1, 2, 4, 8]). This serial design enables hierarchical and deep integration of contextual information across scales, effectively bridging the semantic gap between coarse and fine features. The model is trained end-to-end using a boundary-sensitive structural loss function, which combines a weighted binary cross-entropy loss and a weighted intersection over union(IoU) loss to emphasize learning on challenging boundary pixels.Results and Discussions Comprehensive experiments are conducted on five public benchmarks: DUTS-TE, DUT-OMRON, ECSSD, HKU-IS, and PASCAL-S. LSFNet is compared against 12 state-of-the-art methods. Quantitatively, LSFNet achieves superior or highly competitive performance across all standard metrics [S-measure, weighted F-measure, mean absolute error(MAE), max F-measure, and max E-measure]. For instance, on DUTS-TE, it attains an S-measure of 0.926, an F-measure of 0.899, and an MAE of 0.022, outperforming all compared methods. Notably, it achieves the highest F-measure(0.860) on PASCALS, a dataset known for small and boundary-ambiguous objects, demonstrating its strength in detail preservation. The precision-recall curves and F-measure curves show that LSFNet(red solid line) consistently dominates others, especially maintaining high precision at high recall rates, indicating robust foreground-background discrimination and boundary accuracy. Qualitative visual comparisons further confirm LSFNet’s effectiveness. It produces more precise and structurally coherent saliency maps in challenging cases involving large objects, multiple disconnected objects, small targets, and scenes with low contrast or fine structures. Ablation studies systematically validate the contribution of each proposed module. The full model integrating PDE, EJM, and CFA achieves the best results, proving their synergistic effect. Experiments on the CFA sub-modules confirm the effectiveness of both the implicit enhancement mechanism and the SRPP structure. Furthermore, the design of exponentially increasing dilation rates in SRPP(e.g., [1,2,4,8]) is shown to be more effective than constant or linearly increasing rates. Efficiency analysis shows that LSFNet [18.31×10~9 multiply accumulate operations(MACs), 94.21×10~6 parameters] offers a favorable accuracy-efficiency trade-off compared to models with similar performance but much higher computational costs(e.g., ICON, BBRF). The model’s limitation is also discussed, where it occasionally fails to segment extremely elongated, large-scale objects that span the entire image, a challenge shared with other SOTA methods. This is attributed to insufficient modeling of global continuity for such extreme structures and their scarcity in training data.Conclusions This paper proposes LSFNet, a novel salient object detection network that effectively coordinates global structure and local details. The core innovations include a progressive dual-scale encoder for semantically consistent multi-scale feature extraction, an edge-guided joint modulation module for unsupervised boundary enhancement, and a cross-scale feature aggregation module with serial dilated convolutions for deep hierarchical feature fusion. Extensive experimental results on five benchmarks demonstrate that LSFNet achieves state-of-the-art performance, excelling in both structural integrity and detail richness. Future work will focus on improving the model’s efficiency and its capability to handle objects with extreme aspect ratios or sizes.
【Key words】 image processing; salient object detection; Laplace operator; feature pyramid; cross-scale feature fusion;
- 【文献出处】 激光与光电子学进展 ,Laser & Optoelectronics Progress , 编辑部邮箱 ,2026年12期
- 【分类号】TP391.41
- 【下载频次】12