节点文献

基于XML的PDF文档信息抽取系统的研究

Research on PDF Documents Information Extraction System Based on XML

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 宋艳娟张文德

【Author】 Song Yanjuan (College of Mathematics and Computer Science,Fuzhou Uninversity,Fuzhou 350002,China) Zhang Wende (Library of Fuzhou Uninversity,Fuzhou 350002,China)

【机构】 福州大学数学与计算机科学学院福州大学图书馆 福州350002福州350002

【摘要】 首先设计了科技论文的DTD文档,然后分析了PDF文档的结构。在此基础上,我们介绍了PDF文档信息抽取系统的设计框架。该框架以上述DTD为模板,把以PDF格式表示的科技论文解析转换为有效的XML文档。

【Abstract】 The article is structured as follows.Firstly, we try to design a DTD of articles of science and technology.Secondly,we analyze the structure of PDF documents.Based on that,we dwell on the design of a PDF information extraction system,which use the above-mentioned DTD as a template,transfer a PDF-formatted scientific and technological article to a valid XML document.

【关键词】 信息抽取PDFXML
【Key words】 Information Extraction PDF XML
【基金】 福建省高等学校科技项目(JA04164)的研究成果之一
  • 【文献出处】 现代图书情报技术 ,New Technology of Library and Information Service , 编辑部邮箱 ,2005年09期
  • 【分类号】TP391.1
  • 【被引频次】36
  • 【下载频次】663
节点文献中: 

本文链接的文献网络图示:

本文的引文网络