节点文献
基于KNN及相关链接的中文网页分类研究
Research on Chinese webpages classification based on k-nearest neighbour algorithm and relative hyperlinks
【摘要】 提出了中文网页相关链接提取算法,能够较好地抽取出中文网页中的相关链接,算法的时间复杂性低,准确率和召回率都令人满意.基于向量空间模型,采用KNN对中文网页进行分类,比较了基于网页标题分类、基于网页正文分类,以及将正文与相关链接结合分类、将标题与相关链接结合分类的分类效果,印证了中文网页中相关链接对网页分类具有积极影响的设想,最终分类的准确率达到80%以上.
【Abstract】 This paper proposed an algorithm based on blocking the webpage’s links to retrieve the relative links with good precision,the complexity of the algorithm has the character of time low,and precision and recall are satisfactory.Based on the vector space model,this paper used KNN to classify the chinese webpage.Compared with the results of classification based on the title,classification based on text classification,as well as text and relative links classification together,title and relative links classification together.It was true that the relative links is helpful to the classification of webpage,and the final classification precision is higher than 80%.
【Key words】 chinese webpages classification; webpage theme extraction; relative hyperlinks; KNN;
- 【文献出处】 哈尔滨商业大学学报(自然科学版) ,Journal of Harbin University of Commerce(Natural Sciences Edition) , 编辑部邮箱 ,2011年02期
- 【分类号】TP393.092
- 【被引频次】3
- 【下载频次】118