节点文献

基于聚类算法的网页语义结构分析

ANALYSIS OF WEBPAGE SEMANTIC STRUCTURE BASED ON CLUSTERING

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 梁宏昊邵志清孙晓星

【Author】 Liang Honghao Shao Zhiqing Sun Xiaoxing(Department of Computer Science and Engineering,East China University of Science and Technology,Shanghai 200237,China)

【机构】 华东理工大学计算机科学与工程系

【摘要】 随着语义网的不断发展,网页语义的研究也在不断的进步。但现阶段的网络结构中,非语义化网页仍旧占据了信息系统最主要的部分。信息系统在整合的过程中,也需要了解网页的语义结构以完成信息的获取和分析。提出一种基于视觉特征筛选的网页语义结构分析方法。该方法可以在忽略网页语义的情况下,通过网页结构的视觉特性和内容特性分析网页中不同结构的语义关系,使用聚类分析方法来推定网页中半结构化信息的语义结构,并通过该方法对一组随机网页进行了分析,结果证明该方法具有比较好的分析能力。

【Abstract】 The research on webpage semantics is making constant progress along with the development of semantic web.However,the non-semantic Web pages are still the principal parts of the information systems at present.In the process of information system integration,there is also the need to understand the semantic structure of the Web pages as to accomplish the access and analysis of the information.This paper proposes an approach for analysing semantic structure of the Web pages based on visual feature selection.In circumstance of ignoring the webpage semantics,the approach can analyse the semantic relations with different structures in Web pages by means of visual and content features of the webpage structure,and infer the semantic structures of the semi-structured information in Web pages by cluster analysis.A series of random Web pages have been analysed by this approach.The result turned out that the approach excels in analysis.

【基金】 国家自然科学基金(61003126);上海市自然科学基金(09ZR1408400)
  • 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2012年03期
  • 【分类号】TP18;TP393.092
  • 【下载频次】116
节点文献中: 

本文链接的文献网络图示:

本文的引文网络