节点文献

专题式Web信息获取技术研究

Research of Topic-Specific Web Resource Discovery

【作者】 欧歌

【导师】 赵恒永;

【作者基本信息】 北京化工大学 , 计算机应用技术, 2005, 硕士

【摘要】 Web信息获取存在已经有十几年的历史,近年来网络信息量飞速增长,使得传统的综合性信息获取的发展变得越来越困难,他无法及时的收集所有信息,而且由于信息数量太多,在准确率上无法满足人们的需要。固此,小型的专题信息采集成为近年的研究热点,具备了极高的研究价值。 本文论述了Web信息获取的用途、历史、现状及发展,介绍了信息获取系统的主要流程,对其中现在比较流行的主要算法进行了介绍和比较,分析了中国目前在化工专业方向的网络信息分布情况。使用Java以及SQL Server 2000数据库构建了一个专题式的Web信息获取系统,其中利用元搜索引擎的原理采用人工加机器的方式从网络上收集种子,通过提供全面、准确的网站网址,简化数据过滤的工作,并且在此基础上实现了高效、灵活的信息下载功能。对在HTML的解析,文件过滤中遇到的问题提出了解决的方法,对整个系统的性能及未来的发展提出了总结。 从最后的结果来看,这套系统的方案是行之有效的,获取到的页面质量很好。相信本课题的研究成果也能够适用于其他方向的专题信息获取。

【Abstract】 Web crawler have exist for many years. The rapid growth of the World-Wide Web poses unprecedented scaling challenges for general-purpose crawlers recently. It can not gather all data timely and it is hard to find out the useful information. So the focused web crawler becomes the focus research. The goal of it is to selectively seek out pages that are relevant to a set of topics. It can improve the crawler’s performance, leads to savings in hardware and network resources.In this paper we introduce the uses, history, actuality and future of the focused web crawler, analyse the popular algorithm and distribution of the pages that are relevant to a topic in the web. Build a focused crawler with Java and SQL Server 2000.Collect seeds from web based on metasearch engine theory. Simplify the information filtering through providing comprehensive and exact URL of web site and realize the high effective information crawling. We also give the solution to problems met in analyzing HTML syntax and file filtering. Finally, we make a summary of the capability and the future of the system.The experiment result show that the work is effective and our

【关键词】 信息获取专题搜索引擎种子
【Key words】 Web Resource GatheringTopicSearch EngineSeed
  • 【分类号】TP393.092
  • 【被引频次】1
  • 【下载频次】209
节点文献中: