节点文献
Web文本信息抽取与挖掘方法
Methods for Information Extraction and Data Mining from Web
【摘要】 Web信息资源中蕴含着具有巨大潜在价值的知识。人们迫切需要能够从Web上快速、有效地发现资源和知识的工具。搜索引擎不能完全满足这一要求,为此需要开发比信息检索层次更高的新技术。文中以Web文本为研究对象,着眼于将数据挖掘技术应用于Web挖掘。兼顾中英文文本,提出了一套Web文本的特征表示、特征提取及Web页面的结构化转换方法,并将粗糙集理论应用于转换后的Web文本挖掘。
【Abstract】 A tool, better than information retrieval and can promptly pick up the required knowledge from the vast volume of information in Web is researched. The Web text is considered to be the research object. Aiming to use data mining techniques on Web mining, methods for the Web text feature expression, feature selection and hypertext conversion into a structured database are proposed. These methods can be applied to both Chinese and English Web. Also, the rough set theory is used to mine the converted Web text.
【关键词】 Web挖掘;
特征提取;
Web文本结构化;
粗糙集;
【Key words】 Web mining; feature selection; hypertext conversion; rough set;
【Key words】 Web mining; feature selection; hypertext conversion; rough set;
【基金】 吉林省教育厅资助项目(吉教合字1999第06号)
- 【文献出处】 长春工业大学学报(自然科学版) , 编辑部邮箱 ,2002年S1期
- 【分类号】TP393.092
- 【被引频次】27
- 【下载频次】471