节点文献

骨骼—角度—曲率多模态融合的遮挡场景手势识别

Occlusion scene gesture recognition using skeleton-angle-curvature multimodal fusion

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王琪崔荣一赵亚慧

【Author】 WANG Qi;CUI Rongyi;ZHAO Yahui;School of Information Technology, Changchun Finance College;School of Engineering, Yanbian University;

【机构】 长春金融高等专科学校信息技术学院延边大学工学院

【摘要】 针对手势识别在单目视觉下面临遮挡干扰与精度不足的问题,提出一种融合二维骨骼关键点、手关节角度与手指曲率多模态数据的手势识别方法。为实现数据融合,构建了基于弯曲传感器数据手套与单目相机的采集系统,收集了7名受试者手势动作(包含遮挡场景)的数据集。通过提取手部关键点并计算关节角度,将骨骼、关节角度与曲率信息拼接为多模态输入,进而利用卷积神经网络—双向长短期记忆(CNN-BiLSTM)混合网络分别学习空间与时间特征。实验结果表明,所提出的多模态融合方法相比仅使用骨骼信息,识别准确率从68.34%显著提升至84.13%,证明融合手指曲率与关节角度能有效克服遮挡问题,提高手势识别的鲁棒性与准确性。

【Abstract】 To address the issues of occlusion interference and low accuracy in gesture recognition under monocular vision, a gesture recognition method that integrates multimodal data from 2D skeletal keypoints, hand joint angles, and finger curvature is proposed. To achieve data fusion, a data acquisition system based on a bending sensor data glove and a monocular camera is constructed, and a dataset of gesture actions(including occlusion scene)is collected from seven subjects. By extracting hand keypoints and calculating joint angles, the skeletal, joint angle, and curvature information are concatenated into a multimodal input. Then, a convolutional neural network-bidirectional long short-term memory(CNN-BiLSTM) hybrid network is used to learn spatial and temporal features, respectively. Experimental results show that the proposed multimodal fusion method significantly improves the recognition accuracy from 68.34 % to 84.13 % compared to using only skeletal information, demonstrating that fusing finger curvature and joint angles effectively overcomes the occlusion problems and improves the robustness and accuracy of gesture recognition.

【基金】 2025年吉林省职业教育与成人教育教学改革研究重点课题(2025ZCZ024);2024年长春金融高等专科学校科研规划项目(2024JZ013)
  • 【文献出处】 传感器与微系统 ,Transducer and Microsystem Technologies , 编辑部邮箱 ,2026年06期
  • 【分类号】TP212;TP391.41
  • 【下载频次】35
节点文献中: 

本文链接的文献网络图示:

本文的引文网络