节点文献

基于Transformer的证件图像无检测文字识别

Non-detection text recognition of certificate image based on Transformer

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 肖慧辉张东波王旺王家奎

【Author】 XIAO Hui-hui;ZHANG Dong-bo;WANG Wang;WANG Jia-kui;School of Automation and Electronic Information, Xiangtan University;Wuhan Veilytech Co.,Ltd.;

【通讯作者】 张东波;

【机构】 湘潭大学自动化与电子信息学院武汉唯理科技有限公司

【摘要】 深度学习在图像识别的现存模型中,都有检测和识别两个过程,且需借助复杂的网络结构、大量的文本框标注来提高识别准确率。文中针对存在的问题提出了一个简单且鲁棒性强的证件图片无检测文字识别方法,通过嵌入二维特征图中不同序列位置的水平、竖直方向位置编码,将不同子空间的特征表达连接到序列解码器,解码器部分加入了全局上下文模块,网络模型能并行训练并可以快速收敛,通过插入特殊符号直接得到结构化的字段,简化了信息后处理流程,单张图片识别时间在122ms左右。测试结果表明,模型在身份证扫描件文本图像识别上表现出优越的性能。

【Abstract】 The existing models of deep learning in image recognition have two steps, including detection and recognition, and use the complex network structure and a large number of bounding box annotations to improve the recognition accuracy. The paper proposes a simple and robust method of non-detection text recognition for certificate image. This method directly embeds the horizontal and vertical position coding of different sequence positions in the two-dimensional feature map, and connects the feature expression of different subspaces to the sequence decoder. The decoder part is added with the global context module. The network model can be trained in parallel and can be converged quickly. Moreover, the structured text field can be readily obtained by inserting special symbols, which simplifies the information post-processing process. The recognition time of a single image is about 122 ms. The test results show that the model has excellent performance in text image recognition of scanned ID card.

  • 【文献出处】 信息技术 ,Information Technology , 编辑部邮箱 ,2021年06期
  • 【分类号】TP391.41
  • 【下载频次】455
节点文献中: 

本文链接的文献网络图示:

本文的引文网络