节点文献
基于XML的PDF文档信息抽取系统的研究
Research on PDF Documents Information Extraction System Based on XML
【摘要】 首先设计了科技论文的DTD文档,然后分析了PDF文档的结构。在此基础上,我们介绍了PDF文档信息抽取系统的设计框架。该框架以上述DTD为模板,把以PDF格式表示的科技论文解析转换为有效的XML文档。
【Abstract】 The article is structured as follows.Firstly, we try to design a DTD of articles of science and technology.Secondly,we analyze the structure of PDF documents.Based on that,we dwell on the design of a PDF information extraction system,which use the above-mentioned DTD as a template,transfer a PDF-formatted scientific and technological article to a valid XML document.
【基金】 福建省高等学校科技项目(JA04164)的研究成果之一
- 【文献出处】 现代图书情报技术 ,New Technology of Library and Information Service , 编辑部邮箱 ,2005年09期
- 【分类号】TP391.1
- 【被引频次】36
- 【下载频次】663