节点文献

基于页面Block的Web档案采集和存储

Collecting and Storing Web Archive Based on Page Block

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 宋杰王大玲鲍玉斌申德荣

【Author】 SONG Jie , WANG Da-Ling, BAO Yu-Bin, SHEN De-Rong (School of Information Science and Engineering, Northeastern University, Shenyang 110004, China)

【机构】 东北大学信息科学与工程学院东北大学信息科学与工程学院 辽宁沈阳100004辽宁沈阳100004

【摘要】 提出了基于页面Block对Web页面的采集和存储方式,并详细表述了该方法如何完成基于布局页面分区、Block主题的抽取、版本和差异的比较以及增量存储的方式.实现了一个Web归档原型系统,并对所提出的算法进行了详细的测试.理论和实验表明,所提出的基于页面Block的Web档案(Web archive)采集和存储方法能够很好地适应Web档案的管理方式,并对基于Web档案的查询、搜索、知识发现和数据挖掘等应用提供有利的数据资源.

【Abstract】 In this paper, the page block based Web archive collecting and storing approach is proposed. The algorithms of layout-based page partition, extracting topic from block, version comparison and incremental storage implementation are introduced in detail. The prototype system is implemented and tested to verify the proposed approach. Theoretics and experiments show that, the proposed approach adapts the Web archive management well, and provides a valuable data resource to the Web archive based query, search, data mining and knowledge discovering applications.

【关键词】 Web档案页面分区页块
【Key words】 Web archivepage partition, page block
【基金】 Supported by the National Natural Science Foundation of China under Grant Nos.60573090, 60673139 (国家自然科学基金)
  • 【文献出处】 软件学报 ,Journal of Software , 编辑部邮箱 ,2008年02期
  • 【分类号】TP311.10
  • 【被引频次】47
  • 【下载频次】466
节点文献中: 

本文链接的文献网络图示:

本文的引文网络