节点文献

基于标记树表示方法的页面结构分析

Web Page Structure Analysis Based on Tag Tree Method

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 常育红姜哲朱小燕

【Author】 Chang Yuhong 1,2,3 Jiang Zhe 2,3 Zhu Xiaoyan 2,31 (Beijing Jiuzhou Company,Beijing100081) 2 (Department of Computer Science and Technology,Tsinghua University,Beijing100084) 3 (State Key Laboratory of Intelligent Technology and Systems ,Tsinghua University,Beijing100084)

【机构】 北京九州公司清华大学计算机科学与技术系清华大学计算机科学与技术系 北京100081 清华大学计算机科学与技术系北京100084 清华大学智能技术与系统国家重点实验室北京100084北京100084

【摘要】 页面内容结构分析在WEB信息检索、分类和抽取等方面有重要作用。文章从页面布局和内容之间关系出发,根据WEB文件中标记之间关系,用标记树表示页面文件,采用自底向上的算法,抽取出具有不同语义的页面内容,提出用树形层次结构表示它们之间关系的方法。在此基础上,通过模仿人们浏览页面的习惯,成功地将其应用于页面的计算机屏读系统,实现自动朗读页面主题的功能。

【Abstract】 WEB page content structure is very helpful for applications such as information retrieval,classification,information extraction etc.This paper analyzes the structure of WEB page according to the relation between layout and content.This paper uses tag tree to denote the WEB content and presents a down-top approach to extract the different semantic contents of page,and brings forward a tree structure to present the relation of them.At last it applies it successfully in screen reader for reading WEB pages according to simulating the habit of person browsing WEB page.

  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2004年16期
  • 【分类号】TP393.092
  • 【被引频次】82
  • 【下载频次】379
节点文献中: 

本文链接的文献网络图示:

本文的引文网络