节点文献

基于语料的Web页面抽取器的研究与实现

Research and Implementation of Web Page Extrator Based on Corpus

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陆剑江钱培德

【Author】 LU Jianjiang,QIAN Peide (Dept.of Computer Science and Engineering,Southeast Universisy,Nanjing 210096)

【机构】 东南大学计算机科学与工程系东南大学计算机科学与工程系 南京 210096南京 210096

【摘要】 主要介绍了面对万维网上各种各样的诸如文本、声音、图形和图像等语料信息,如何按照用户的实际需求将其中对用户有用的信息抽取出来,从而实现对现有语料信息的一种有效分离。重点介绍了Web信息簇聚性的特点和语料库的设计,以及语料库的实际工作原理。

【Abstract】 This thesis mainly discusses how to extract the useful information of corpus according to the user’s actual requirement from the World Wide Web where there are all kinds of information of corpus such as text,sound,image and picture,etc.By using this method,people can realize the useful extraction from the current existing information of corpus. It emphases the fascination specialty of information in the World Wide Web and the actual working principle of the database of corpus.

【关键词】 Web语料超文本标记语言可扩展标记语言
【Key words】 WebCorpusHTMLXML
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2003年06期
  • 【分类号】TP393.092
  • 【被引频次】19
  • 【下载频次】105
节点文献中: 

本文链接的文献网络图示:

本文的引文网络