节点文献

网页分类技术

Web document classification techniques

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙建涛沈抖陆玉昌石纯一

【Author】 SUN Jiantao, SHEN Dou, LU Yuchang, SHI Chunyi Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China)

【机构】 清华大学计算机科学与技术系智能技术与系统国家重点实验室清华大学计算机科学与技术系智能技术与系统国家重点实验室 北京100084北京100084北京100084

【摘要】 网页分类是使用机器学习的方法实现网页类别的自动标注。回顾了文本分类技术的研究状况,分析了网页的结构特征,难点在于结合网页的结构信息选择合理的表示方式和分类算法。使用纯文本分类技术处理网页是不合理的。基于概率模型的方法和关系学习方法计算量大,关系学习方法学习结果的可解释性好,支持向量机方法分类准确率高,但核函数的构造和大规模数据集的训练都是该算法的难题。应该采用多种指标对网页分类算法进行评价。

【Abstract】 Web document classification assigns labels to web documents based on machine learning techniques. A review of various text classification techniques showed that the main difficulties in web document classification are the page representation methods and the classification algorithms. Techniques that go beyond text categorization approaches are needed. Probabilistic algorithms and relational learning methods are both time-consuming. SVM (support vector machine) classifiers are quite accurate but the automatic kernel selection and the large scale training are both key problems. Various measures were investigated to compare algorithm performance based on sample datasets.

【基金】 国家"九七三"基础研究基金项目(G1998030414)
  • 【文献出处】 清华大学学报(自然科学版) ,Journal of Tsinghua University(Science and Technology) , 编辑部邮箱 ,2004年01期
  • 【分类号】TP393.092
  • 【被引频次】94
  • 【下载频次】1272
节点文献中: 

本文链接的文献网络图示:

本文的引文网络