节点文献
3D CNN人体动作识别中的特征组合优选
Feature combination optimizing in 3D CNN human motion recognition
【摘要】 为了提高人体动作识别准确率,改进原有3D CNN网络模型以获得更为丰富细致的人体动作特征,并通过对比实验为模型输入优选出识别效果最好的特征组合。该模型主要包括5个卷积层、3个下采样层和2个全连接层,二次卷积操作有利于提取到更为细致的特征,BN算法和dropout层用以防止模型过拟合,空间金字塔池化技术可以使网络能够处理任何分辨率的图像,提高模型适用性。通过在KTH和UCF101数据集上做识别测试实验,特征组合"ViBe二值图+光流图+三帧差分图"作为模型输入可以得到较高的识别准确率,尤其针对背景较复杂、动作类别多且差异性较小的数据集提高明显,具有较好的实际应用价值。
【Abstract】 In order to improve the accuracy of human motion recognition, a new 3 D CNN network model is constructed to obtain more detailed human motion features, and the best combination of features is selected through comparative experiments as input of the model. The model consists of five convolution layers, three undersampling layers and two full connection layers. The secondary convolution operation is beneficial to extract more detailed human motion features, BN algorithm and dropout layer are used to prevent model over-fitting. Spatial pyramid pooling technology can enable the network to process any resolution image and improve the applicability of the model. Through the recognition test on KTH and UCF101 data sets, the combination of feature "vibe binary graph + optical flow graph + three frame difference map" as model input can obtain higher recognition accuracy, especially for the data set of complex background, multiple action categories and small differences which has obviously improved and has good practical application value.
【Key words】 deep learning; human motion recognition; three-dimensional convolution neural network; BN algorithm; dropout technology; spatial pyramid pooling;
- 【文献出处】 河北工业大学学报 ,Journal of Hebei University of Technology , 编辑部邮箱 ,2021年01期
- 【分类号】TP391.41;TP183
- 【下载频次】138