节点文献
回转库档案盒标签图像的检测识别研究
【作者】 李婧;
【作者基本信息】 南京师范大学 , 电子与通信工程(专业学位), 2021, 硕士
【摘要】 档案作为最原始的一手资料具有极其重要的保存价值。纸质档案通常装订成册放入盒中进行保管。档案室大量的、高密度的档案盒存储特点使其日常管理工作存在效率低下、人工成本高等问题。随着现代化信息技术的不断发展,档案的管理也越来越趋向于数字化。如何利用现代信息技术手段实现档案盒的自动检测识别以达到高效管理档案的目的是诸多科技工作者一直致力于研究解决的问题。本文针对一种空间利用率较高的回转库档案柜的存储特点,研究了基于档案盒标签图像的检测识别技术,通过OCR、图像匹配等数字图像处理技术识别比对档案盒标签,该技术能够大大提高档案盒的盘点管理效率。论文的主要工作及相关创新点如下:(1)分析比较了传统图像分割算法及深度学习分割算法对档案盒标签图像分割的优劣。传统数字图像处理方法采用经典的边缘检测分割算法以及基于阈值的分割算法,深度学习分割算法采用分割性能极佳的U-net神经网络。实验结果显示传统分割算法受噪声影响较大,许多不定参量需要根据应用场景调节,无法做到完全自适应;而采用U-net神经网络对档案盒标签图像进行分割,没有任何参数需要调节,且抗干扰能力强,分割效果好。(2)研究了采用Tesseract-OCR文字识别技术进行档案盒标签的检测识别方法。为了能够对档案盒标签“竖排”的文字进行识别,论文研究采用投影法分割标签内文字个体,并讨论了使用图像形态学处理方法增强标签的分割效果;为提高字符识别正确率,针对档案盒标签的字体进行了字体库数据增强训练,训练结果较好地提高了字符识别效果。(3)研究了SURF算法进行图像匹配的档案盒标签检测技术,并通过各种干扰测试了SURF算法的稳定性。针对实际环境下的档案盒标签图像的特点进行了相应的图像增强处理,包括自动色彩均衡算法,去噪算法等,提升了实际环境下档案盒标签图像的匹配正确率。(4)为了进一步提高SURF匹配算法的匹配准确率,并根据回转库档案盒标签图像的颜色及摆放位置的特点,提出了若干个改进的SURF匹配算法,即:基于HSV颜色提取的SURF匹配算法、基于RANSAC算法的SURF匹配算法以及基于双向匹配及距离优化的SURF匹配算法。这些改进的SURF匹配算法都在一定程度上提高了匹配正确率。
【Abstract】 As the most original first-hand information,archives have extremely important preservation,which are usually bound into volumes and put into boxes for safekeeping.There are a large number and high density of file boxes that exist problems of low efficiency and high labor costs in its daily inventory.With the continuous development of modern information technology,the management of archives tends to be increasingly digital.Thus,how to utilize modern information technology to realize the automatic detection and identification of archival boxes to achieve efficient archival management is a problem that many scientific and technical workers have been devoted to addressing.Given the storage characteristics of a rotating storage cabinet with relatively high space utilization,this paper studies the detection technology based on the image of the file box label and identifies and compares the file box label through OCR,image matching and other digital image processing technology,which can greatly improve the efficiency of the inventory of the file box.The main work and innovation of this paper are as follows:(1)The advantages and disadvantages of traditional image segmentation algorithms and deep learning segmentation algorithms are analyzed and compared.Traditional digital image processing methods use classical edge detection segmentation algorithms and threshold-based segmentation algorithm.Deep learning segmentation algorithm uses a U-net neural network with excellent segmentation performance.The experimental results show that the traditional segmentation algorithm is greatly affected by noise,and plenty of uncertain parameters need to be adjusted according to the application scene,which can not be fully adaptive;while using the U-net neural network to segment the file box label image is much better since there is no need to adjust any parameters and has the strong anti-interference ability and good segmentation effect.(2)It mainly focuses on the detection and recognition method of file box labels by using Tesseract-OCR character recognition technology.To realize the character recognition of the "vertical row" of the file box label,it studies the use of the projection method to segment the individual text in the label,and discusses the use of image morphology processing method to enhance the segmentation effect of the label.For improving the accuracy of character recognition,the font database data of the file box label is enhanced,and the training results can better enhance the character recognition effect.(3)This paper studies the file box label detection technology of the SURF algorithm for image matching and tests the stability of the SURF algorithm through various interferences.In the light of the file box label image in the actual environment corresponding image enhancement processing,including automatic color equalization algorithm,denoising algorithm,promote the actual environment of the file box label image matching accuracy.(4)Aiming to further improve the matching accuracy of the SURF matching algorithm,the SURF matching algorithm based on HSV color extraction,the SURF matching algorithm based on RANSAC algorithm,and the SURF matching algorithm based on two-way matching and distance optimization were proposed according to the characteristics of the color of the label images and placement of the box of the rotating library.These improved SURF algorithms all upgrade the matching accuracy to a certain extent.
【Key words】 Image segmentation; Character recognition; SURF matching algorithm;