节点文献

基于模糊分类的印刷体数学公式抽取方法

Mathematical formula extraction method from printed document based on fuzzy classification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 田学东郝楠

【Author】 TIAN Xue-dong,HAO Nan (College of Mathematics and Computer Science,Hebei University,Baoding Hebei 071002,China)

【机构】 河北大学数学与计算机学院河北大学数学与计算机学院 河北保定071002河北保定071002

【摘要】 公式抽取是印刷体数学公式识别的基础性环节,现有的识别方法多以公式区域已知为前提,相关的研究还很欠缺。通过引入模糊分类理论,提出了一种孤立数学公式的抽取算法,通过对大量训练样张的数据统计与分析,选取了非规则度、宽高比、密度等6维特征,由此构建出对孤立公式行、文本行、标题行的模糊分类规则,实现了孤立公式行的抽取。实验结果表明,该方法有较高的准确性和鲁棒性。

【Abstract】 Process of mathematical formula extraction from printed document is a basal step.Most of the available extraction methods assume that the regions containing mathematical formulas are known.An algorithm to extract isolated mathematical formulas by introducing fuzzy classification theory was described.Six features,such as degree of irregularity,width-to-height ratio and density ect,were selected from lots of data that came from training samples counted and analyzed,thereby the rule of fuzzy classification was built to handle isolated mathematical formula lines,text lines and title lines,so mathematical formula extraction was realized.The experimental results indicate that this method could obtain favorable veracity and good robustness.

【基金】 河北省科学技术研究与发展计划资助项目(06213598)
  • 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2007年08期
  • 【分类号】TP391.4
  • 【被引频次】10
  • 【下载频次】103
节点文献中: 

本文链接的文献网络图示:

本文的引文网络