节点文献

一种基于频繁路径特征的XML文档结构聚类算法改进实现

An Improved XML Document Structural Clustering Algorithm Using Frequent Path Patterns

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 姚俊杰崔斌

【机构】 南开大学计算机科学与技术系北京大学计算机科学系

【摘要】 <正>1引言随着XML数据的持续增加,有效处理XML数据并为决策提供信息支持变得日益重要。XML允许半结构化和层次化表示,对XML数据的挖掘不同于传统结构化数据和文本数据。XML的挖掘研

【Abstract】 With the standardization of XML as an information exchange language over the Web,a huge amount of information is formatted in XML documents.The semi-structural nature of XML means that,the structural information is an important clue as to the meaning of XML documents.The structural clustering result plays an important role in many web information retrieval and XML manipulation tasks.As to the short coming of a path-based XML document clustering algorithm on robust and scalability,we have the algorithm improved by replacing similarity measure,weighting clustering features and introducing new algorithms.The experiments show that our improved algorithm outperforms the old one and is scalable well.

【基金】 国家自然科学基金项目(60603045)
  • 【会议录名称】 第二十四届中国数据库学术会议论文集(技术报告篇)
  • 【会议名称】第二十四届中国数据库学术会议
  • 【会议时间】2007-10-20
  • 【会议地点】中国海南海口
  • 【分类号】TP311.10
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络