节点文献

基于深度学习的视觉运动放大方法及实验研究

Research on Visual Motion Magnification Based on Deep Learning

【作者】 陈力

【导师】 彭聪;

【作者基本信息】 南京航空航天大学 , 控制科学与工程, 2023, 硕士

【摘要】 现实世界中充满微小运动,它们难以被肉眼察觉,却能揭示世界的内在机制和变化趋势。运动放大是微小运动研究的重要手段。传统运动放大方法存在处理效率低、运动感知能力弱、噪声特性差的问题。近年来,计算机视觉与人工智能迅速发展,有望为运动放大研究带来突破。本文基于深度学习开展运动放大研究,构建运动放大数据集,提出多种深度学习运动放大方法,并开展详尽的实验研究。首先,针对深度学习运动放大的数据缺失问题,构建运动放大合成数据集和真实数据集,二者分别支持监督学习和无监督学习。合成数据集通过图像剪贴和仿射变换模拟逼真的纹理和运动。该数据集包含运动放大真值,以支持监督学习运动放大方法。真实数据集基于光流选取公共数据集中的真实图像,以支持无监督学习运动放大方法。其次,针对拉格朗日方法处理效率低、欧拉方法运动感知能力差的问题,提出一种基于监督学习的多任务运动放大方法。该方法通过多任务学习融合拉格朗日和欧拉两类主流运动放大方法,构建多任务网络,以实现优势互补。该网络根据欧拉框架构建,取得高效运动放大过程,并从拉格朗日方法引入光流,获取准确运动感知,实现高效优质的运动放大。与现有方法开展对比实验,验证该方法具备优越的放大性能、可靠的鲁棒性以及出众的处理效率。再次,针对监督学习运动放大方法生成图像真实性低、复杂运动处理效果差的问题,提出一种基于无监督学习的运动放大方法。该方法改进监督学习运动放大方法的网络结构,并基于运动放大处理的可重复性设计重编码策略,规避真值缺失,以实现无监督的运动放大。该网络根据高效的欧拉框架构建,基于真实图像按照重编码策略进行训练,实现逼真的运动放大效果。与现有方法开展对比实验,验证该方法具备一定的放大性能和较好的生成图像真实性。最后,针对提出的深度学习运动放大方法,构建出视觉运动放大系统。该系统由软件平台和硬件设备构成,支持实时运动放大和视频运动放大的工作模式,集成传统方法和深度学习方法在内的多种运动放大技术。开展实时和视频运动放大实验,验证该系统能在各个工作模式下可视化微小运动,为运动放大技术的应用提供通路。

【Abstract】 The world is full of tiny motions and deformations.These variations are nearly invisible to naked eyes but reveal internal mechanisms and variations tendencies of the world.Motion magnification is an effective tool to visualize tiny motions and deformations.Conventional approaches either require complex computation or cannot distinguish subtle motions from noise.Computer vision and artificial intelligence technologies have developed rapidly in recent years,which is expected to bring a breakthrough to motion magnification.This paper conducts motion magnification research via deep learning.We establish motion magnification datasets and propose several deep learning motion magnification approaches.The provided experiments verify the proposed approaches in detail.Firstly,deep learning motion magnification approaches relies on abundant data,synthetic and real datasets thus are established to support supervised and unsupervised learning.The synthetic dataset sticks foregrounds on backgrounds to generate images and imitates motions by affine transformations.The real dataset selects real images from public datasets according to optical flow.Secondly,conventional approaches either require complicated computation or cannot distinguish tiny variations from noise.To counter such problem,a multi-task motion magnification approach based on supervised learning is proposed.The approach fuses Lagrangian and Eulerian methods,two mainstream solutions to motion magnification,to establish a multi-task network for mutual advantage complement.The network is mainly developed from Eulerian methods for efficient inference and introduce optical flow from Lagrangian methods for precise motion perception.The proposed approach is evaluated by qualitative and quantitative experiments,revealing promising performance,reliable robustness,and high efficiency.Thirdly,supervised learning approaches cannot generate realistic images and is weak in processing complicated motions.To handle such problem,an unsupervised learning motion magnification approach is proposed.The approach improves network architectures of the supervised learning method and designs a re-encode strategy in accordance with the repeatability of motion magnification to avoid the lack of real magnified images.The network is developed from efficient Eulerian methods and trained on real images by the re-encode strategy for realistic motion magnification.The proposed approach is evaluated by qualitative and quantitative experiments,presenting acceptable performance and superior image authenticity.Finally,a visual motion magnification system is established for the proposed deep learning motion magnification approaches.The system is composed of a software platform and hardware equipment.The system supports real-time and off-line motion magnification modes and integrates several motion magnification approaches.The experiments verifies that the system can visualize tiny motions in different modes.

  • 【分类号】TP18;TP391.41
节点文献中: