节点文献

文本页面中数学表达式的定位及分析

Orientation and Analysis of Mathematical Expressions in Document Images

【作者】 陈波

【导师】 王加俊;

【作者基本信息】 苏州大学 , 通信与信息系统, 2007, 硕士

【摘要】 电子文档具有容易修改、检索和传输等优点,从而基于移动办公终端的文档实时电子化变得越来越频繁。文档的电子化必须经过页面分割和字符识别,页面内通常含有多种元素如字符、图片、表格和数学表达式等,其中数学表达式的分析、识别和重组是文档电子化的难点。因此研究高效的分析算法十分必要,本文的工作主要体现在以下几个方面:鉴于文本页面各文本元区域的前景像素存在自相关性,本文提出了基于微结构的页面分割算法来切分文本页面。首先采用快速扫描算法将前景像素归类并形成微结构集,利用微结构的相关性分类出页面含有的图元、表格元等;改变合并规则合并分类后的字符元得到字符区,选取字符区域的最大者结合最小二乘法检测字符区的倾斜角度来校正页面;最后利用微结构并结合水平投影将校正后的页面切割为文本行。数学表达式的二维结构特性使数学表达式行与普通文本行存在很大差异,本文利用这些差异将独立表达式行与普通文本行区分开来;接着采用连通体搜索方法搜索分类后的文本行,判断搜索得到的连通体与该文本行上下基线的关系确定内嵌表达式所在位置,结合最大投影间隔法切分出内嵌表达式,最后借助微结构和投影法分析数学表达式结构。实验结果表明,本文提出的算法是有效的,并具有较好的稳定性、适应性。此外,将文本元逐个分类和分解会增加识别的成功率,更加有利于字符的识别。

【Abstract】 Since the electronic document has advantages such as its convenience in revision, retrieval and transmission, the real-time transformation of the traditional paper document to its electronic version based on the mobile office terminal becomes more and more frequent. Two processes of document image segmentation and character recognition are needed to realize the transform. Usually, many elements such as words, images, graphics, forms and mathematical expressions are contained in document images. Difficulties of the electronic transformation and document reusing lie in the analyzing, recognizing and rebuilding of the mathematical expressions. To tackle the above problem, it is necessary to develop efficient algorithms for the electronic transformation. Contributions of this thesis are as follows:Owing to the existence of the self-correlations among foreground pixels in different regions of the document images, a page segmentation algorithm based on the microstructures is proposed in this thesis. Firstly, the foreground pixels in the document image are classified into different microstructure sets with a fast scanning algorithm, and the elements such as the halftones and forms are classified with the correlations between the microstructures. Then, the rest of the character structures are merged by changed rules. The skew angle of the document image is detected from the largest merged character region with the least square algorithm, after which the skew correction is performed on the document image with the detected skew angle. Finally, the de-skewed document image is segmented into text lines by means of the horizontal project profile in combination with the microstructure.There are large differences between ordinary text lines and the lines containing mathematical expressions because of the two-dimensional structure of the expressions. In this thesis, the independent expressions lines are separated from the ordinary text lines according to the above differences. Then the in line expressions contained in the classified text lines are located according to the relationships between the connected components and the upper and lower base lines, after which, the in line expressions are segmented with the method of the maximum projection gap. Lastly, the mathematical expressions are analyzed with the microstructure and the projecting method. Experimental results show that the proposed algorithm is effective and also has superiorities of stability and flexibility. In addition, the most prominent superiority of this method lies in that the document elements are classified and decomposed one by one through structural classification, which can improve the recognition performance. Therefore, the OCR for the decomposed elements will also benefit from it.

  • 【网络出版投稿人】 苏州大学
  • 【网络出版年期】2008年 04期
  • 【分类号】TP391.1
  • 【被引频次】1
  • 【下载频次】90
节点文献中: 

本文链接的文献网络图示:

本文的引文网络