节点文献

网络科技信息监测中富文档识别与信息提取技术研究

Identification and Information Extraction of Rich Documents for Web Scientific Information Monitoring

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张敏刘建华谢靖

【Author】 ZHANG Min;LIU Jian-hua;XIE Jing;National Science Library,Chinese Academy of Sciences;University of Chinese Academy of Sciences;

【机构】 中国科学院文献情报中心中国科学院大学

【摘要】 【目的/意义】围绕富文档载体类型的鉴别、元数据的提取等开展相应的实际应用探索。【方法/过程】通过开源工具PDFBox以及Tika对不同类型的富文档元数据及正文内容进行提取,取得了良好的实际效果,为科研人员提供了大量的有学术价值的情报资源。【结果/结论】通过对富文档监测与识别的研究与探索,笔者拓展了文本知识内容的识别方法,为后续的深度知识分析提供了有效的支撑。

【Abstract】 【Purpose/significance】This paper focuses on the practical application of the identification of the rich documentcarrier, the extraction of metadata and the content of the text, and so on.【Method/process】Through the open source tools,such as PDFBox and Tika, the author provides a lot of valuable information resources for the scientific research personnel,which has obtained good actual effect.【Result/conclusion】With the survey and identification of rich documents, the authorexpands the identification methods of text knowledge contents,and provides the effective support to the coming deep knowl-edge analysis.

【基金】 中国科学院文献情报能力建设专项(院1509);教育部人文社科基金(14YJC870029)
  • 【文献出处】 情报科学 ,Information Science , 编辑部邮箱 ,2017年01期
  • 【分类号】G254
  • 【被引频次】7
  • 【下载频次】180
节点文献中: 

本文链接的文献网络图示:

本文的引文网络