节点文献

文档识别中误切分字符拒识问题的研究

Research on the Missegmented Character Rejection in Document Recognition

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈臻刚丁晓青刘长松彭良瑞

【Author】 Chen Zhengang Ding Xiaoqing Liu Changsong Peng Liangrui (State Key Laboratory of Intelligent Techniques and Systems ,Department of Electronic Engineering,Tsinghua University,Beijing100084)

【机构】 清华大学电子工程系智能技术与系统国家重点实验室清华大学电子工程系智能技术与系统国家重点实验室 北京100084北京100084北京100084

【摘要】 自动文档识别中字切分算法如果仅仅依靠大小位置等度量信息,很容易产生误切分图像块,需要字符分类器给出一定的反馈才能准确切分,为此提出了一个新的拒识算法,目标是尽可能准确地拒识非法字符。该文分析了基于距离的分类器的置信度和广义置信度,在此基础上改进了常用的广义置信度映射函数,并设计了一个基于样本学习的拒识规则,提高了拒识算法的适应性。在中日韩三种文档样本上的实验表明,该文算法明显改善了系统性能,对于较低质量的印刷文本识别具有一定的普遍意义。

【Abstract】 In OCR systems the character segmentation algorithm may generate missegmented blocks,especially when us-ing only geometric measure information such as size and location.Feedback information from character classifier is nec-essary to achieve higher character segmentation accuracy.In this paper a novel rejection algorithm is proposed to reject these invalid characters more accurately.First,the confidence and generalized confidence of distance-based classifiers are analyzed,and then usual generalized confidence mapping function is modified.A new sample-based rejection rule is also proposed,which is more adaptive and flexible.Experiments on Chinese,Japanese and Korean document recognition show that new rejection algorithm evidently improved the system performance,especially for low-quality printed document recognition.

【关键词】 OCR字符识别置信度拒识规则
【Key words】 OCRCharacter RecognitionConfidenceRejection Rule
【基金】 国家863高技术研究发展计划(编号:2001AA114081);国家自然科学基金(编号:69972024)
  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2002年17期
  • 【分类号】TP391.43
  • 【被引频次】16
  • 【下载频次】195
节点文献中: 

本文链接的文献网络图示:

本文的引文网络