节点文献

面向场景文本识别的语义独立深度学习方法研究

Research on Semantics Independence Deep Learning Methods for Scene Text Recognition

【作者】 王彬;

【导师】 许信顺;

【作者基本信息】 山东大学 , 软件工程, 2022, 硕士

【摘要】 文字自诞生起就承载着传递信息的责任。随着互联网技术与移动技术的迅速发展,存储在计算机上的信息以指数级别爆炸增长。文本作为信息的载体也因此增长迅速。对于计算机而言,存储与利用自然场景中的信息是以红绿蓝为基础元素的图片。令计算机自动识别图片中的文本信息有着广泛的应用意义,如自动驾驶、票据识别、人机交互等。近几年来,效果最好的模型大都是基于视觉语义的模型,这类方法通常先使用一个特征提取器将二维的图片提取为视觉特征图,接着利用语义模型(或称为语言模型)对前面得到的特征图进一步编码,得到语义特征,然后综合利用视觉特征与语义特征得到最终的识别结果。但是这种方式后一步的语义模型通常高度依赖视觉特征。这种将语义特征与视觉的特征耦合的方式有两个缺点:一是语义模型多是沦为视觉模型的纠正器,仅仅用于更正视觉模型得到的结果。语义模型在整个流程端对端训练,但是其实际意义却是作为后处理部分,这就导致了模型的冗余,梯度链增长难以训练。其次利用语言模型对于视觉模型的结果进行纠正,这种方式的确可以很大的提高准确率,但是自然场景应用中广泛存在着错误的文本信息,比方说手写试卷识别与批改,对于错误的文本,模型会识别后自动将其纠正为正确的,这大大偏离了批改的本意。为解决上述提出的问题,本文提出了一种新颖的语义独立网络(Semantics Independence Network),将语义模块独立出来,使之成为与视觉模型对等的部分,使得视觉模型更加关注二维的视觉特征,而语义模型更加关注一维的语义特征。此外,本文提出视觉语义融合模块,将视觉特征与语义特征充分交互。通过以上两种方式,语义模块可以独立地处理语义信息,视觉与语义模块可以充分解耦,并且又充分利用了两部分的特征。针对目前文本识别网络冗余的问题,本文提出用于分析场景文本识别模型模块参数冗余的剪枝方法,为设计场景文本识别网络时是否使用某一模块提供了检验方式,本文提出冗余参数修剪的方式并引入了层感知的剪枝率设置,对本文提出的语义独立的场景文本识别方法进行分模块的后剪枝,有针对性地分析了本文提出的语义模块,融合模块的有效性以及当下场景文本识别网络广泛使用Transformer网络参数的冗余性。本文的主要贡献如下:(1)提出了一种语义独立的文本识别方法,它不同于之前模型仅仅利用截断梯度来解耦视觉与语义模型,而是从模型结构上进行调整,实现结构上彻底的解耦。(2)提出了一种新的视觉语义特征融合模块,摆脱了二维视觉特征与一维语义特征之间的语义鸿沟问题,充分利用了视觉特征与语义特征。(3)提出了一种用于冗余参数修剪的方式并将其用于本文识别模型,针对本文提出的语义独立模型场景文本识别方法的各个模块进行剪枝,并分析了各个模块的冗余程度。(4)引入了层感知的剪枝率设置,并将其运用于本文提出的冗余参数修剪方法,将不同待剪层之间的差异考虑进剪枝方法,有针对性的对不同层进行剪枝。

【Abstract】 In recent years,with the rapid development of Internet and mobile technologies,the amount of information stored on computers has exploded exponentially.As a carrier of information,the number of scene text images is also growing rapidly.Information is stored and utilized by computers in natural scenes through two-dimensional images based on red,green and blue.Automatic recognition of text information in images by computers has a wide range of applications,such as automatic driving,bill recognition,human-computer interaction,etc.In recent years,the models which have achieved the best results are mostly based on vision and semantics.These methods usually begin by using a feature extractor to extract visual features from two-dimensional images.Then the semantics model(also known as the language model)is used to encode the feature maps of the previous step to obtain the semantic features.Finally,the visual and semantic features are used to obtain the final recognition result.In this way,semantics models are highly dependent on visual features.This method of coupling semantic features with visual features has two disadvantages:Firstly,semantics models become correctors for visual models,which are only used to correct the results obtained by vision models.Secondly,using semantics models to correct the results of vision model can greatly improve the accuracy.However,there is a wide range of wrong text information in natural scene applications,such as handwritten text recognition and marking.For incorrect text,the model recognizes and automatically corrects it,which is a departure from the purpose of our mission.Moreover,due to the serial coupling of vision and semantics model,the model is bloated and difficult to train.In order to solve the above problems,a novel Semantics Independence Network is proposed in this thesis,which separates semantics modules.The semantics module is separated and becomes the equivalent part of the vision model,so that the vision model pays more attention to the two-dimensional visual features and the semantics model pays more attention to the one-dimensional semantic features.In addition,a vision semantics fusion module is proposed to fully interact visual features with semantic features.Through the above two ways,the semantics module can process semantic information independently.Vision and semantics module can be fully decoupled,and the features of the two parts can be fully utilized.In addition,a pruning method for analyzing the parameter redundancy of the scene text recognition model is proposed for the first time.It provides a way to check whether a module should be used when designing a scene text recognition network.After pruning the trained model by modules,the number of parameters is effectively reduced and the function of each module is verified.A redundant parameter pruning method is proposed and layer-aware pruning rate setting is introduced.Through the post-pruning of the proposed semantics independence scene text recognition method,the validity of the proposed semantics module and fusion module and the redundancy of Transformer network parameters are analyzed.The main contributions of this thesis are summarized as follows:(1)A text recognition network based on semantics independence is proposed,which is different from the previous model that decoupled vision model and semantics model by truncating gradient.Instead,it adjusts the structure of the model to achieve complete decoupled structure.(2)A new visual features and semantic features fusion module is designed,which makes the visual features and semantic features fully interact and make full use of the visual features and semantic features.(3)A new pruning method for redundant parameters is designed and is applied to the text recognition network.It prunes each module of the text method which we proposed above and analyzes the redundancy of each module.(4)A layer-aware pruning rate setting is introduced into the pruning method mentioned above.By considering the difference between different layers to be cut into the pruning method,different layers are pruned differently.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2023年 02期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络