节点文献
空间时间增强Transformer的三维人体姿态估计
Spatial-temporal enhanced Transformer for 3D human pose estimation
【摘要】 由于航母舰载机起降时的噪音和各种设备的干扰,甲板上指挥员的语音指令难以清晰传达。因此,肢体动作作为有效的沟通方式,提高了指挥效率。为检验指挥员动作的安全性和规范性,需要对指挥员进行姿态估计。现有Transformer方法能够有效捕捉视频中关节间的全局依赖关系,但在精确建模局部依赖关系方面仍存在局限性。为同时关注关节的全局与局部信息,提出一种基于空间时间增强Transformer(STEFormer)的三维人体姿态估计方法。该方法通过交替使用空间增强Transformer块和时间增强Transformer块,分别学习关节间运动的空间相关性和各关节运动的时间相关性。此外,为进一步融合全局与局部信息,提出一种增强多头自注意力模块(EMSA),该模块结合了可变维度多头自注意力和卷积神经网络。为提升模型的估计精度和鲁棒性,采用了包含3个误差项的损失函数作为优化目标。在2个数据集上进行实验验证,所提方法在Human3.6M数据集上将MPJPE和P-MPJPE指标分别降低至39.4 mm和31.3 mm,在MPI-INF-3DHP数据集上PCK、AUC和MPJPE指标分别达到99.3%、88.0%和14.6 mm。实验结果表明,STEFormer方法能有效提升人体姿态估计的精度,尤其在减少人体四肢末端关节的估计误差方面表现更为显著,并展现出较强的泛化能力。
【Abstract】 On aircraft carriers, noise from aircraft operations and interference from various equipment make it difficult to clearly convey voice commands from deck directors. Body gestures therefore serve as an effective means of communication, improving command efficiency. To ensure the safety and standardization of these gestures, pose estimation of the director is required. Existing Transformer-based methods can effectively capture global dependencies among joints in videos, but still exhibit limitations in accurately modeling local dependencies. To simultaneously address both global and local joint information, a Spatial-temporal Enhanced Transformer(STEFormer) is proposed for 3D human pose estimation. The method alternates between spatial-enhanced and temporal-enhanced Transformer blocks to learn spatial correlations across joints and temporal dependencies of individual joint movements, respectively. Furthermore, an Enhanced Multi-head Self-Attention(EMSA) module is introduced to better integrate global and local information, combining variable-dimensional multi-head self-attention with a convolutional neural network. To enhance estimation accuracy and robustness, a loss function incorporating three error terms is adopted as the optimization objective. Experiments on two benchmark datasets validate the effectiveness of the proposed approach. On the Human3.6M dataset, MPJPE and P-MPJPE are reduced to 39.4 mm and 31.3 mm, respectively. On the MPI-INF-3DHP dataset, it achieves a PCK of 99.3%, an AUC of 88.0%, and an MPJPE of 14.6 mm. The results demonstrate that STEFormer significantly improves human pose estimation accuracy, particularly in reducing errors for distal limb joints, and exhibits strong generalization capability.
【Key words】 commander; 3D human pose estimation; spatial-temporal enhanced transformer; enhanced multi-head self-attention; convolution neural network; loss function;
- 【文献出处】 兵器装备工程学报 ,Journal of Ordnance Equipment Engineering , 编辑部邮箱 ,2026年04期
- 【分类号】TP391.41;TP18
- 【下载频次】22