节点文献
一种以主题词为特征的排序划界网站聚类算法
Websites Clustering Algorithm through Sorting and Delimiting them Characteristic Value
【摘要】 为了设计出基于自我意识的语义Web中网眼agent的自我意识即Web的语义结构,文中分析了网站及网页的组成元素和组成结构,收集互联网的主题词并精简处理得到Web特征词,根据特征词的出现频率对主题词进行排序;把这个主题词序列作为网站特征值的结构,根据这些特征词在某个网站上是否出现把对应位置为1或0,计算出网站的特征值;提出并实现了一种基于主题词的网站特征值排序划界聚类算法,最后通过两组数据对该算法的有效性进行了验证.
【Abstract】 In order to design the semantic structure of Web-the self consciousness of WOW agent in the semantic Web based on self consciousness,this paper analyzes the constituent elements and structure of Websites and their Web pages.Subject words are collected and simplified to obtain characteristic words of Web.And then the subject words are sorted according to the frequency of the occurrence of the characteristic words.With the sequence of subject words used as the structure of the website characteristic value,according to whether these characteristic words appear on a website or not,the corresponding positions are set as 1 or 0 and the characteristic values of the website are determined.Finally,a clustering algorithm for sorting characteristic value is proposed and its effectiveness is verified by using two sets of data.
【Key words】 self-consciousness; the semantic Web; the semantic structure of Web; subject words; clustering;
- 【文献出处】 西安工业大学学报 ,Journal of Xi’an Technological University , 编辑部邮箱 ,2018年03期
- 【分类号】TP391.1;TP393.092
- 【被引频次】1
- 【下载频次】46