节点文献
网络科技信息监测中富文档识别与信息提取技术研究
Identification and Information Extraction of Rich Documents for Web Scientific Information Monitoring
【摘要】 【目的/意义】围绕富文档载体类型的鉴别、元数据的提取等开展相应的实际应用探索。【方法/过程】通过开源工具PDFBox以及Tika对不同类型的富文档元数据及正文内容进行提取,取得了良好的实际效果,为科研人员提供了大量的有学术价值的情报资源。【结果/结论】通过对富文档监测与识别的研究与探索,笔者拓展了文本知识内容的识别方法,为后续的深度知识分析提供了有效的支撑。
【Abstract】 【Purpose/significance】This paper focuses on the practical application of the identification of the rich documentcarrier, the extraction of metadata and the content of the text, and so on.【Method/process】Through the open source tools,such as PDFBox and Tika, the author provides a lot of valuable information resources for the scientific research personnel,which has obtained good actual effect.【Result/conclusion】With the survey and identification of rich documents, the authorexpands the identification methods of text knowledge contents,and provides the effective support to the coming deep knowl-edge analysis.
【Key words】 rich documents; metadata; identification of the rich document carrier;
- 【文献出处】 情报科学 ,Information Science , 编辑部邮箱 ,2017年01期
- 【分类号】G254
- 【被引频次】7
- 【下载频次】180