节点文献

基于.NET的Web信息抽取系统关键技术研究

The Critical Technology Research on Web Information Extraction System Based on .NET

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 谭锋; 李天真; 崔亮亮;

【Author】 Tan Feng Li Tianzhen Cui Liangliang

【机构】 湖州职业技术学院机电工程分院; 浙江久立集团股份有限公司;

【摘要】 随着Web信息抽取的研究和发展,抽取技术已经逐渐成熟,通过软件来实现从Web页中抽取所需要的信息已成为可能。对基于.NET技术实现的Web信息抽取系统进行了研究,分析并提出了HTML文档下载和清理、HTML到XML格式转换、数据定位及抽取、抽取数据的保存等需要研究解决的关键技术问题,并探讨了相应的解决方案。

【Abstract】 With the Web information extraction researchment and development,and in the extraction technology has gradually matured through the software from a Web page to extract the required information is possible.Based on.NET technology for Web information extraction system for research,analysis and put forward the document to download and clean up HTML,HTML to XML format,data location and extraction,extraction of data preservation needs to study and solve key technical problems and to explore the corresponding solutions.

【关键词】 .NET; Web信息抽取; 应用软件; HTML; XML;
【Key words】 .NET; Web Information Extraction; Application Software; HTML; XML;
【基金】 浙江省教育厅科研项目(Y200803750)
  • 【文献出处】 软件导刊 ,Software Guide , 编辑部邮箱 ,2010年12期
  • 【分类号】TP311.52
  • 【被引频次】2
  • 【下载频次】80
节点文献中: