节点文献
基于频繁子树模式的半结构化数据集聚类
Semi-structured data cluster based on frequent subtree pattern
【摘要】 为提高大数据时代半结构化数据集聚类分析效率,提出一种以数据集频繁子树模式为特征的半结构化数据集聚类方法。提出一种频繁子树模式挖掘方法FSTPMiner,使用“编码树”数据结构对半结构化数据进行编码,通过编码树将树结构频繁模式挖掘过程转化为线性表结构频繁模式挖掘,提高挖掘效率。使用频繁子树模式作为特征并构建特征向量空间,基于经典凝聚型层次聚类方法对半结构化文档数据集进行聚类。经过对照实验,与Costa算法、ICQB算法和Damalagas算法相比,在保证聚类结果正确率前提下,对半结构化数据集聚类效率方面具有优势。
【Abstract】 To improvc the efficiency of clustering analysis of semi-structured data sets,a scmi-structured data clustering method charactcrized by frequent subtree patterns of data sets was proposed.A FSTPMiner posted for mining frequent subtree was proposed,a coding tree CT data structure was proposed to encode scmi-structured data,and the tree structure frequent pattern mining process was transformed into a linear structure frequent pattern mining through the coding tree to improve mining efficiency.Furthermore,the frequent subtree patterns were used as features and the feature vector space was constructed,and the semistructured document data set was clustered based on the classic agglomerated hierarchical clustering method.Compared with Costa algorithm,ICQB algorithm and Damalagas algorithm in the experiment,the method FSTPMincr guarantees the accuracy of clustering results,and the clustering effieieney of semi-structured data is improved.
【Key words】 big data; semi-structured data; frcquent subtrce pattcrn; cluster; coding tree;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2022年10期
- 【分类号】TP311.13
- 【下载频次】52