节点文献

中文电子文档中数学公式的语义识别方法研究

Research on Semantic Recognition for Mathematical Formula in Digital Chinese Documents

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王高王培珍杜培明王爱芳张自强

【Author】 WANG Gao;WANG Pei-zhen;DU Pei-ming;WANG Ai-fang;ZHANG Zi-qiang;School of Electrical Engineering & Information,Anhui University of Technology;

【机构】 安徽工业大学电气与信息工程学院

【摘要】 中文电子文档中数学公式结构复杂且含有大量特殊符号,针对目前OCR技术难以高效识别数学公式,提出了一种新的公式语义识别方法.首先结合字符宽度中心矩和汉字拒识法对公式进行两次定位,然后利用投影法和连通域法切分公式字符,提取字符孔洞数、穿越线等特征构建字符模板库,利用模板匹配方法识别公式中各字符,接着基于五类特征字符的特点,建立后标型、包含型和独立型等七种字符块合并规则以分析公式结构、还原公式的语法含义,最后将公式结构分析结果以EQ域语法串的形式输出.实验结果表明,本文方法可以有效地对中文电子文档中的数学公式进行语义分析.

【Abstract】 As a result of the complex structure and large numbers of special symbols in mathematical formulas,it’s difficult for OCR technology to recognize formulas efficiently from digital Chinese documents at present.In view of this,a novel formula semantic recognition approach is proposed.Firstly,the paper locates formulas twice using both central moment of character width and rejection of Chinese character method.Then,the projection method and connected domain are put forward on character segmentation,and hole number,traversing line and some other features are extracted from characters to create character template library,characters in the formula are recognized by template matching.Next,in order to analyze the structure and grammatical meaning of the formula,seven combination rules are established based on the characteristics of five kinds of characters.Finally,structural analysis results output in the form of EQ domain syntax string.Experimental results show that the proposed method can realize semantic analysis for mathematical formulas in digital Chinese document effectively.

【基金】 国家自然科学基金项目(51574004)资助
  • 【文献出处】 小型微型计算机系统 ,Journal of Chinese Computer Systems , 编辑部邮箱 ,2017年10期
  • 【分类号】TP391.41
  • 【被引频次】1
  • 【下载频次】236
节点文献中: 

本文链接的文献网络图示:

本文的引文网络