节点文献

中文网页分类及存储系统设计与实现

The Design and Implementation of Chinese Webpage Classification and Storage System

【作者】 于成龙

【导师】 苏小红; 闫立辉;

【作者基本信息】 哈尔滨工业大学 , 软件工程, 2007, 硕士

【摘要】 随着互联网技术的快速发展和网络资源的迅速膨胀,为了给用户提供高效准确的服务,我们需要对网络中纷繁复杂的资源进行合理的组织与分类。网络上的信息资源有着海量、动态、异构、半结构化等显著特点,由于缺乏统一的组织和管理而显得杂乱无章,给Web检索带来了一定的困难。使用网页分类技术可以更加有效地组织和管理网络资源,提高信息检索的效率,它目前已成为网络检索的研究热点之一。由于网络具有一定的更新性,网络资源每隔一段时间就会被网站自动的更新掉,这样再去查询这些信息时,用户可能会搜索不到。目前,我国还没有开始意识到进行网络资源永久保存的重要性。永久保存各时期的网站内容,可以防止有用的网络资源被永远更新掉,保护了网络资源,也方便用户查找任一时期的网站信息。本论文以网页分类以及与之相关的网页信息抽取处理、分词、特征提取和增量式存储技术为网页分类和存储手段,通过对中文网页结构和永久性存储的全面分析来构建中文网页分类及存储系统,该系统能够对采集下来的网络资源进行准确分类,并可以进行增量式存储,这样可以便于用户查询,同时也有效的节省了存储空间。本文介绍了中文网页信息抽取处理、分词、特征提取方法,并对中文网页的结构和特点进行了分析,提出中文网页分类和存储的设计与实现方法,最后通过程序设计语言来实现,并进行了测试和验证。测试结果达到了系统设计的要求,应用效果显著。

【Abstract】 Along with Internet technology being fast development and networkresources rapid inflation,in order to provide the highly effective andaccurate service to the user,we need to deal with the complicated resourcesin the network by the reasonable organization and the classification.In thenetwork information resource has the remarkable characteristics of the masscapacity,the dynamic,the isomerism,half structure and so on,as a result,thelacks of the unified organization and management appear chaotically,andhave brought difficulties for the Web retrieval.In the network informationresource has the capacity for alcohol,the tendency,the isomerism, halfstructure and so on.The remarkable characteristic,as a result,lacks of theunification the organizationand the management appear chaotically,andhave brought the certain difficulties for the Web retrieval.The usingwebpage classification technology may effectively organize and managenetwork resources,and enhance the efficiency of retrievaling information,ithas at present become one of hot spots of the network retrieval research.Because the network has the certain renewal,the network resources canbe updated every now and again,when inquiring these informations likethis,the user possibly can not search them.Now,our country has not startedto realize the importance of preserving the network resourcespermanently.The permanently preserved website content of the various timemay prevent the useful network resources renewed forever,and protect thenetwork resources,and also facilitate the user searching the websiteinformation any time.The paper use extracting,segmenting,classifying webpages anddistilling characteristics as the methods of increasedly storing webpageinformation and classifying webpages,through comprehensive analysis tothe structure of Chinese webpage and permanent storage,it constructsChinese webpage classification and storage system.This system canaccurately classify the resources which gather from the network and may carry out the incremental storage,so it may be advantageous for the users’inquiries,simultaneously,effectively has saved the storage space.This article introduced the methods of extracting,segmenting Chinesewebpage information,and distilling the characteristic.It has analysedChinese webpage structures and characteristics,and proposed the design andthe realizing methods of Chinese webpage classification and memorysystem.The test results have achieved the system design requirement,theapplication effect is remarkable.

  • 【分类号】TP391.1
  • 【被引频次】2
  • 【下载频次】293
节点文献中: