节点文献
基于3D卷积神经网络的低照度视频增强技术
Low Illumination Video Enhancement Technology Based on 3D Convolutional Neural Network
【作者】 孟焱;
【导师】 门爱东;
【作者基本信息】 北京邮电大学 , 电子与通信工程(专业学位), 2021, 硕士
【摘要】 随着高性能并行计算芯片的快速更新迭代,很多以往由于规模庞大或模型复杂,仅在理论中出现的神经网络结构逐渐可以得以应用,而3D卷积就是其中一种技术。当前CV领域内研究人员已经能够在视频动作分类,3D图像分割等多种任务中灵活采用3D卷积构建深度神经网络模型,并达到领域最先进水平。在低照度视频增强任务中,主要有基于单帧图片的增强和基于连续多帧图片的增强两个方向。在基于连续多帧图片的方向上,当前已有一些较为理想的算法,但鲜有对3D卷积的应用。本文首先面向基于连续多帧的低照度动态视频增强这一任务,构建了基于3D卷积神经网络的低照度视频增强模型3D-Unet。为了对这一模型进行训练,在一个目标识别数据集上通过筛选、降低亮度、加入噪声等处理生成了一个低照度视频数据集。最后将这一模型与目前最先进的低照度视频增强方法从图像恢复效果和视频恢复效果两个方面对比。最终实验结果表明,将3D卷积应用以3D-Unet模型的形式应用于低照度动态视频恢复任务是可行的,并且性能优于当前最优方法。
【Abstract】 With the rapid updating and iteration of high-performance parallel computing chips,many neural network structures,which were previously only found in theory due to their large scale or complex models,can be applied,and 3D convolution is one of them.At present,researchers in the field of CV have been able to flexibly use 3D convolution to construct deep neural network models in video action classification,3D image segmentation and other tasks,and have reached the most advanced level in the field.In the low illumination video enhancement task,there are two directions:one is based on single frame image enhancement and the other is based on continuous multiframe image enhancement.At present,there are some algorithms based on the direction of continuous multi-frame images,but there are few applications of 3D convolution.In this thesis,a low-illuminance video enhancement model 3 D-UNET based on 3D convolutional neural network is firstly constructed for the task of low-illuminance dynamic video enhancement based on continuous multi-frames.In order to train the model,a low illumination video dataset is generated on a target recognition dataset by filtering,reducing brightness and adding noise.Finally,this model is compared with the most advanced low-illuminance video enhancement methods from two aspects of image recovery and video recovery.The final experimental results show that the application of 3D convolution in the form of 3D-UNET model is feasible for low illumination dynamic video recovery task,and the performance is better than the current optimal method.
- 【网络出版投稿人】 北京邮电大学 【网络出版年期】2024年 01期
- 【分类号】TP391.41;TP183