节点文献
基于模糊分类的印刷体数学公式抽取方法
Mathematical formula extraction method from printed document based on fuzzy classification
【摘要】 公式抽取是印刷体数学公式识别的基础性环节,现有的识别方法多以公式区域已知为前提,相关的研究还很欠缺。通过引入模糊分类理论,提出了一种孤立数学公式的抽取算法,通过对大量训练样张的数据统计与分析,选取了非规则度、宽高比、密度等6维特征,由此构建出对孤立公式行、文本行、标题行的模糊分类规则,实现了孤立公式行的抽取。实验结果表明,该方法有较高的准确性和鲁棒性。
【Abstract】 Process of mathematical formula extraction from printed document is a basal step.Most of the available extraction methods assume that the regions containing mathematical formulas are known.An algorithm to extract isolated mathematical formulas by introducing fuzzy classification theory was described.Six features,such as degree of irregularity,width-to-height ratio and density ect,were selected from lots of data that came from training samples counted and analyzed,thereby the rule of fuzzy classification was built to handle isolated mathematical formula lines,text lines and title lines,so mathematical formula extraction was realized.The experimental results indicate that this method could obtain favorable veracity and good robustness.
【Key words】 printed mathematical formula recognition; formula extraction; fuzzy classification;
- 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2007年08期
- 【分类号】TP391.4
- 【被引频次】10
- 【下载频次】103