节点文献

面向高光谱图像分类的记忆增强型视觉Transformer方法

A Memory-enhanced Vision Transformer Method for Hyperspectral Image Classification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 唐楠张猛陈润涵

【Author】 Tang Nan;Zhang Meng;Chen Runhan;School of Computer, Central China Normal University;School of Mathematics and Physics, China University of Geosciences;

【通讯作者】 张猛;

【机构】 华中师范大学计算机学院中国地质大学(武汉)数学与物理学院

【摘要】 针对高光谱图像分类中Transformer模型自注意力机制计算复杂度高、在高维小样本条件下易产生模型冗余与过拟合的问题,本文提出一种面向高光谱遥感数据的轻量化视觉Transformer方法(Vision Transformer with Attention Memory Integration,Vi TAMin)。该方法通过引入可学习的记忆单元,在网络层间高效传递全局上下文信息,以“令牌-记忆单元”交互机制替代传统的“令牌-令牌”自注意力计算,将计算复杂度显著降低至O(zn)(其中z为记忆单元数量,n为令牌数量,z< n)。此外,本文设计了一种轻量级记忆更新策略,通过自适应融合历史记忆与当前令牌特征,实现跨窗口样本信息的有效建模,增强了模型全局表征能力。在4个公开数据集上的实验结果表明,Vi TAMin的整体分类精度(Overall Accuracy, OA)分别达到99.49%、99.73%、99.28%和99.01%,其它相关指标也均优于主流卷积神经网络(Convolutional Neural Network, CNN)及Transformer方法,在保证高分类精度的同时显著降低了模型的计算与存储开销。实验结果验证了Vi TAMin在复杂高光谱遥感场景下的有效性与鲁棒性,为高维地学数据的高效建模与分析提供了一种新的思路,并具备在地球物理数据处理等相关领域中的应用潜力。

【Abstract】 To address the high computational complexity of the self-attention mechanism in Transformer models for hyperspectral image classification, as well as the issues of model redundancy and overfitting under high-dimensional and small-sample conditions, we propose a lightweight vision Transformer tailored for hyperspectral remote sensing data, termed Vi TAMin(vision Transformer with attention memory integration). This method introduces learnable memory units to efficiently propagate global contextual information across network layers. By replacing the conventional “token-to-token” self-attention with a “tokento-memory” interaction mechanism, the computational complexity is significantly reduced to(where is the number of memory units, and denotes the number of tokens, with). In addition, a lightweight memory update strategy is designed to adaptively fuse historical memory with current token features, enabling effective modeling of cross-window sample information and enhancing the global representation capability of the model. Experimental results on four public datasets demonstrate that Vi TAMin achieves overall classification accuracies(OA) of 99.49%, 99.73%, 99.28%, and 99.01%, respectively. Other evaluation metrics also consistently outperform mainstream CNN(convolutional neural network)-and Transformerbased methods. While maintaining high classification accuracy, the proposed method substantially reduces both computational and storage costs. These results validate the effectiveness and robustness of Vi TAMin in complex hyperspectral remote sensing scenarios, providing a novel approach for the efficient modeling and analysis of high-dimensional geoscientific data, with promising potential for applications in related fields such as geophysical data processing.

【基金】 国家自然科学基金(编号:42274172);重庆市自然科学基金(编号:2023NSCQ-MSX0207)
  • 【文献出处】 工程地球物理学报 ,Chinese Journal of Engineering Geophysics , 编辑部邮箱 ,2026年03期
  • 【分类号】TP751;TP18
  • 【下载频次】10
节点文献中: 

本文链接的文献网络图示:

本文的引文网络