节点文献

多模态注意力感知与相邻尺度建模的Transformer网络

Transformer Network with Multimodal Attention Perception and Adjacent-Scale Modeling

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 宋霄罡张浩泽张小龙赵钦黑新宏何敏

【Author】 SONG Xiaogang;ZHANG Haoze;ZHANG Xiaolong;ZHAO Qin;HEI Xinhong;HE Min;School of Computer Science and Engineering, Xi’an University of Technology;Human Machine Integration Intelligent Robots Shaanxi Provincial University Engineering Research Center, Xi’an University of Technology;School of Civil Engineering and Architecture, Xi’an University of Technology;

【通讯作者】 何敏;

【机构】 西安理工大学计算机科学与工程学院西安理工大学人机共融智能机器人陕西省高校工程研究中心西安理工大学土木工程与建筑学院

【摘要】 RGB-D显著目标检测旨在从配对的彩色图像与深度图像中识别最具视觉吸引力的目标,关键问题在于实现多模态特征与多尺度特征的有效融合.现有方法在RGB特征与深度特征融合过程中,模态互补信息表达、边缘细节保持及尺度间关联利用仍有进一步提升空间.为此,文中提出多模态注意力感知与相邻尺度建模的Transformer网络(Transformer Network with Multimodal Attention Perception and Adjacent-Scale Modeling, MATNet).首先,采用双分支金字塔池化Transformer编码器,分别提取RGB模态和深度模态的多层级特征,并在各层引入多模态注意力融合模块,联合通道注意力与空间注意力,增强模态互补信息表达和关键区域语义一致性.然后,构建相邻尺度建模模块,自上而下逐级聚合相邻尺度特征,有效融合高层语义信息与低层边缘纹理信息,提升显著目标的结构完整性与边界表征能力.最后,结合多尺度预测与监督机制,构建端到端检测框架.在5个公开数据集上的实验表明,MATNet在提升检测精度与边缘保持能力方面具有稳定性与有效性.

【Abstract】 RGB-D salient object detection aims to identify the most visually attractive objects from paired color images and depth images, and the key challenge is the effective fusion of multimodal and multiscale features. The existing methods still need the improvement in modal complementary information representation, edge detail preservation and utilization of cross-scale association during the fusion of RGB features and depth features. Therefore, a Transformer network with multimodal attention perception and adjacent-scale modeling(MATNet) is proposed. Multilevel RGB features and depth features are extracted by dual-branch pyramid pooling Transformer encoders. A multimodal attention fusion module is introduced at each stage. The modal complementary information representation and semantic consistency in key regions are jointly enhanced by channel attention and spatial attention. Then, an adjacent-scale modeling module is constructed to aggregate adjacent-scale features progressively in a top-down manner. High-level semantic information and low-level edge texture information are fused effectively. The structural integrity and boundary representation capability of salient objects are improved. Finally, an end-to-end detection framework is constructed by combining multi-scale prediction and the supervision mechanism. Experiments on five public datasets demonstrate that MATNet is effective and stable in improving detection accuracy and edge preservation capability.

【基金】 国家自然科学基金联合基金重点项目(No.U2568225);国家自然科学基金项目(No.52372418,U2368203);陕西省创新能力支持计划项目(No.2025RS-CXTD-006)资助~~
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2026年04期
  • 【分类号】TP391.41;TP183
  • 【下载频次】22
节点文献中: 

本文链接的文献网络图示:

本文的引文网络