节点文献

基于深度学习的实时轻量型人像分割方法研究

Research on Real-time Lightweight Portrait Segmentation Method Based on Deep Learning

【作者】 吴迪;

【导师】 黄东晋;

【作者基本信息】 上海大学 , 计算机科学与技术, 2023, 硕士

【摘要】 人像分割是通过对图像中每个像素进行分类,将图像划分为人像和背景。由于人像和背景的结构复杂多样,高效、准确的从自然背景中分割出人像面临着巨大的挑战。传统人像分割方法大多数只关注图像的颜色和自然结构,分割精度普遍较低。随着计算机算力的提升和深度学习技术的发展,基于深度学习的人像分割技术取得了显著的突破,但是大部分模型仍需要经过复杂的计算无法达到实时的速率。针对上述问题,本文对轻量型网络在复杂背景下的单人像和多人像的实时分割任务进行探究,主要的研究内容和创新工作有:(1)针对自然背景下单人物图像分割不精确的问题,本文提出了基于可变形卷积和跳跃连接的实时单人像分割网络。首先,本文提出了可变形深度可分离卷积块,将可变形卷积与深度可分离的卷积结合,使得网络在获取特征信息时能有效控制时间消耗。其次,通过对跳跃连接中的补充信息单独设置类别置信度阈值,进一步提高了网络分割精度。最后,提出了一种新的损失函数,有效地提高自然背景中人像分割的鲁棒性。实验表明,该网络具有0.122M参数量和0.092G浮点数计算量,能以极轻的模型量级实现自然背景下的单人像实时精确分割,在公共数据集EG1800、人像Matting数据集和会议视频数据集上准确率分别为95.60%、97.63%和97.36%,取得了最优的综合性能。(2)针对自然背景下多人物形态多变,互相遮挡等问题,本文提出了基于多层空洞卷积和信息融合的实时多人像分割网络。通过在网络的深层放置了多个感受野,以增强网络获取信息的能力。为了控制模型复杂度,本文使用空洞卷积和分组卷积进行了绝大部分的特征提取操作,且对空洞卷积的扩张系数进行了超参数优化,有效提升了模型精度。另外,通过采用像素重排列和空间注意力机制对多尺度信息进行融合,保证了信息融合的有效性。实验表明,该网络具有0.049M参数量和0.164G浮点数计算量,能以极轻的量级实现自然背景下的多人像实时精确分割,在多人像公共数据集MHP v1.0(Multi-Human-Parsing)上准确率为86.70%,取得了最优的综合性能。

【Abstract】 Portrait segmentation is the process of classifying each pixel in an image and dividing it into portrait and background.Due to the complex and diverse structures of portraits and backgrounds,efficient and accurate segmentation of portraits from natural backgrounds faces enormous challenges.Most traditional portrait segmentation methods only focus on the color and natural structure of the image,and the segmentation accuracy is generally poor.With the improvement of computer computing power and the development of deep learning technology,significant breakthroughs have been made in portrait segmentation technology based on deep learning.However,most models still require complex calculations and cannot achieve real-time speed.In response to the above issues,this article explores the real-time segmentation task of single person and multiple person images in complex backgrounds for lightweight networks.The main research content and innovative work include:(1)This dissertation proposes a real-time single portrait segmentation network based on deformable convolution and skip connections to address the issue of inaccurate segmentation of single person images in natural backgrounds.Firstly,this article proposes a deformable deep separable convolution block,which combines deformable convolution with deep separable convolution,enabling the network to effectively control time consumption when obtaining feature information.Secondly,by setting a separate category confidence threshold for the supplementary information in the skip connection,the network segmentation accuracy is further improved.Finally,a new loss function is proposed to effectively improve the robustness of human image segmentation in natural background.Eexperiments show that the network has 0.122 M parameters and 0.092 G floating point computation,and can achieve real-time accurate segmentation of single person images in natural background at an extremely light model level.The accuracy rates are 95.60%,97.63%and 97.36% respectively on the public dataset EG1800,Portrait Matting Dataset and Conference Video Dataset,achieving the best comprehensive performance.(2)This dissertation proposes a real-time multi portrait segmentation network based on multi-layer hollow convolution and information fusion to address the issues of multiple characters’ varying shapes and mutual occlusion in natural backgrounds.By placing multiple receptive field in the deep layer of the network,the ability of the network to obtain information is enhanced.In order to control the complexity of the model,this dissertation uses atrous convolution and groups convolution for most feature extraction operations,and hyperparameter optimization is carried out for the expansion coefficient of void convolution,which effectively improves the accuracy of the model.In addition,the effectiveness of information fusion is ensured by using pixel rearrangement and spatial attention mechanism to fuse multi-scale information.Experiments show that the network has 0.049 M parameters and 0.164 G floating point computation,and can achieve real-time accurate segmentation of multiple portraits in natural background at an extremely light model level.The accuracy rate on the Multi Human Parsing(MHP v1.0)Dataset is 86.70%,and the optimal comprehensive performance is achieved.

  • 【网络出版投稿人】 上海大学
  • 【网络出版年期】2025年 04期
  • 【分类号】TP391.41;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络