节点文献

基于模板的网页主题信息抽取

Webpage Topic Information Extraction Based on the Template

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 冯少卿都云程施水才

【Author】 FENG Shao-qing~1 DU Yun-cheng~1 SHI Shui-Cai~1 (1.Chinese Information Processing Research Center,Beijing Information Science and Technology University,Beijing 100101,China)

【机构】 北京信息科技大学中文信息处理研究中心

【摘要】 快速准确地抽取网页主题信息是影响 Web 应用服务质量的关键。网页模板就是已经做好的网页框架,由模板生成的网页结构布局是基本一致的。本文提出了利用模板技术进行网页主题信息抽取的算法。该方法充分考虑了网页的结构特征,能够明显改善信息抽取的性能。实验结果表明,该方法准确率可达99.6%。

【Abstract】 Extracting topic information of web pages accurately and efficiently is a key technique to improve the service qualities of web applications.Web pages’structures created by the same template are very identical,because the template provides their framework.A new method to extract topic information based on template was proposed in this paper.The method took much advantage of the characteristic of web page’s structure and improved extraction performance obviously. Experiments indicate that the accuracy reaches 99.6%.

【关键词】 DOM网页样本集模板信息抽取
【Key words】 DOMwebpagesample collectiontemplateinformation extraction
【基金】 863计划重点项目(2006AA010105);北京市属市管高校人才强教计划项目(PXM2007_014224_044677,PXM2007_014224_044676);北京市教委科技发展计划项目(KM200710772010)
  • 【会议录名称】 第三届全国信息检索与内容安全学术会议论文集
  • 【会议名称】第三届全国信息检索与内容安全学术会议
  • 【会议时间】2007-11
  • 【会议地点】中国江苏苏州
  • 【分类号】TP391.1
  • 【主办单位】中国中文信息学会信息检索与内容安全专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络