节点文献
基于模板的网页主题信息抽取
Webpage Topic Information Extraction Based on the Template
【Author】 FENG Shao-qing~1 DU Yun-cheng~1 SHI Shui-Cai~1 (1.Chinese Information Processing Research Center,Beijing Information Science and Technology University,Beijing 100101,China)
【机构】 北京信息科技大学中文信息处理研究中心;
【摘要】 快速准确地抽取网页主题信息是影响 Web 应用服务质量的关键。网页模板就是已经做好的网页框架,由模板生成的网页结构布局是基本一致的。本文提出了利用模板技术进行网页主题信息抽取的算法。该方法充分考虑了网页的结构特征,能够明显改善信息抽取的性能。实验结果表明,该方法准确率可达99.6%。
【Abstract】 Extracting topic information of web pages accurately and efficiently is a key technique to improve the service qualities of web applications.Web pages’structures created by the same template are very identical,because the template provides their framework.A new method to extract topic information based on template was proposed in this paper.The method took much advantage of the characteristic of web page’s structure and improved extraction performance obviously. Experiments indicate that the accuracy reaches 99.6%.
【Key words】 DOM; webpage; sample collection; template; information extraction;
- 【会议录名称】 第三届全国信息检索与内容安全学术会议论文集
- 【会议名称】第三届全国信息检索与内容安全学术会议
- 【会议时间】2007-11
- 【会议地点】中国江苏苏州
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会信息检索与内容安全专业委员会