节点文献
基于规则的HTML文档元数据提取
Metadata Extracting for HTML Document Based on Rules
【摘要】 提出了一种基于规则提取HTML文档元数据的方法,介绍了规则的语法、语义和规则库的设计,研制了一个原型系统MEDES(MEtaData Extracting System),实现HTML文档元数据的自动提取。文章的最后给出了实验结果和评价,并指出进一步的工作。
【Abstract】 This paper proposes a metadata extracting method for HTML document based on rules. After introducing syntax and semantics of the rules, design of rule library is discussed. Based on the method aforementioned, a system named MEDES(metadata extracting system) is developed, which can perform automatic metadata extraction from HTML documents. The experiment results are evaluated, and the future work are discussed in the end.
【关键词】 元数据提取;
基于规则;
信息检索;
Web;
【Key words】 Metadata extracting; Rule controling; Information retrieval; Web;
【Key words】 Metadata extracting; Rule controling; Information retrieval; Web;
- 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2004年09期
- 【分类号】TP393.09
- 【被引频次】17
- 【下载频次】234