节点文献

Java XML与面向Web的智能数据抽取

Intelligence Data Extraction Based on Java XML and Web

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 文艺刘循

【Author】 WEN Yi, LIU Xun (College of Computer Science,Sichuan University, Chengdu 610065,China)

【机构】 四川大学计算机学院四川大学计算机学院 成都610065成都610065

【摘要】 采用标准Web技术———HTML,XML和Java,开发一种基于Web用Java把Web数据转换为XML的数据挖掘方法.该方法标识数据源并把它映射成XHTML,根据一定的相关关系查找数据内的引用点并进行智能数据抽取,将数据映射成XML.这种数据抽取方法比较简单,通过选择可靠的数据源以及在这些数据源中选取与内容相关但与格式无关的锚点,可以较为方便地建立一个强壮的数据抽取系统.

【Abstract】 A method for web-based data mining is developed using the standard technologies of the web--HTML,XML, and Java. convert existing web pages into XML with XML. The data extraction method is very simple only by selecting some reliable data resources and anchor-points which are dependent on those data resources and content of web pages, but independent of the form of web pages.

【关键词】 XMLXHTMLXSL数据抽取
【Key words】 XMLXHTMLXSLdata extraction
  • 【文献出处】 四川大学学报(自然科学版) ,Journal of Sichuan University (Natural Science Edition) , 编辑部邮箱 ,2004年02期
  • 【分类号】TP393.09
  • 【被引频次】14
  • 【下载频次】193
节点文献中: 

本文链接的文献网络图示:

本文的引文网络