节点文献
城市计算中细粒度时空大数据挖掘算法的研究
Fine-Grained Spatiotemporal Data Mining in Urban Computing
【作者】 刘瑞强;
【导师】 程渤;
【作者基本信息】 北京邮电大学 , 计算机科学与技术, 2022, 博士
【摘要】 借助于快速发展的传感技术和大数据技术,城市空间正持续地产生包括交通、多媒体、环境等多个领域下多源异构的海量数据。这些数据除了本身含有的信息外还会带有数据产生时所处的时间与地点信息,具有时空属性,故被称为时空大数据。这些多源异构的时空大数据中蕴含着丰富的知识。对这些时空大数据进行数据挖掘可以帮助理解城市中的各种现象,为解决城市中面临的各种问题带来极大的助力。目前时空大数据挖掘算法已在城市计算中的多个领域得到了广泛的关注,并取得了许多有价值的研究成果。然而当前大部分的时空数据挖掘相关的工作还停留在时空粒度相对较粗的设定下,使得这些工作无法直接应用于细粒度的时空任务,产生的社会价值也相对有限。从粗粒度扩展为细粒度的时空数据挖掘算法仍有许多需要提升的空间,具体包括:(1)城市计算是一个综合性问题,引入跨域数据来协助分析是很有必要的。但现有的跨域数据融合方法都是通过后融合的方式而忽略了跨域数据在相近时隙下的联合特征,这种不充分的融合方式直接导致预测性能低下;(2)细粒度时空设定会使稀疏的分类标签更不平衡,表现为零膨胀问题,这种极稀疏的非零标签会使神经网络难以生效;(3)细粒度时空设定会加剧数据缺失的问题,带有缺失值的数据集也会极大影响算法对数据分布特征的提取,使最后的预测结果变得不准确。针对上述挑战,本文研究了城市计算中由粗粒度扩展到细粒度时空设定下的数据挖掘算法的关键技术,通过深度学习模型强大的高阶特征提取能力以及与城市拓扑空间十分契合的图神经网络算法来对这些时空大数据进行数据挖掘。本文将主要从“如何有效融合跨域数据以支持细粒度的时空数据预测”、“如何对细粒度时空设定而加剧的零膨胀数据进行分析预测”以及“如何对细粒度时空设定加剧的缺失数据进行准确的数据补全”三个方面展开研究。本文的主要内容及贡献如下:1.针对如何有效融合跨域数据以支持细粒度的时空数据预测问题,本文提出了基于跨域数据融合的细粒度时空数据预测算法。该算法模型可以有效地收集跨域数据的时空特征与跨域数据之间的联合特征。与此前基于后融合的算法不同,本文提出的跨域融合算法在模型的前期通过跨域卷积层可以关注到跨域数据在相近时间间隙下的联合特征。本文通过城市计算中一个典型的需要多域特征融合技术的城市异常检测场景对算法进行了验证,并与现有的跨域融合算法进行比较。实验结果显示了该方法的有效性,在异常检测问题中可以取得更高的准确率和召回率。2.针对如何对细粒度时空设定而加剧的零膨胀数据进行分析预测问题,本文提出了基于离散标签连续化策略的稀疏时空数据预测算法,并在上述跨域融合模型的基础上加入了注意力机制,使模型可以在数据极度稀疏的条件下完成时空数据的分析预测任务。提出的离散数据连续化标签增强策略通过对数据进行线性变换和取对数的操作,使原离散数据转换为连续数据,同时把标签元素所对应空间节点历史上的非零元素信息注入到连续数据中。此外注意力机制也可以使模型更好地关注到稀疏时空数据的特征。该算法在事件极度稀疏的城市犯罪预测场景进行了验证。实验结果表明这种离散数据连续化算法可大大改善在分类标签不平衡情况下神经网络模型的性能,与基于注意力机制的融合模型结合可以完成稀疏数据标签条件下的时空数据分析预测。3.针对如何对细粒度时空设定加剧的缺失数据进行准确的数据补全问题,本文提出了基于元学习图注意力机制的细粒度时空数据补全算法。该算法通过挖掘数据的时空相关性来对缺失数据进行补全。模型首先通过图嵌入技术与元学习结合来提取节点之间潜在的空间关系。并结合先验知识加入辅助任务来对补全数据结果进行约束与修正。本文通过智慧交通应用中的流量数据缺失场景与现有的数据补全算法进行了比较。实验中使用了在上海和深圳两个城市采集的真实数据集以及基于多种场景而设置的仿真数据集。实验结果表明本文提出的算法对数据补全的结果更准确,且在不同的场景下都表现出了充分的鲁棒性。
【Abstract】 Thanks to sensing and big data technologies,urban spaces are continuously generating huge amounts of data in many areas such as traffic,multimedia and environment.These data have spatio-temporal properties,i.e.in addition to the information they contain,they also carry information about the time and place in which they were created,hence the term spatio-temporal big data.There is a wealth of knowledge in these multi-source heterogeneous spatio-temporal big data.Data mining of these spatio-temporal data can help to understand the problems faced by cities and can be a great help in solving them.Spatio-temporal big data mining algorithms have received extensive attention in a number of areas of urban computing,and many valuable research results have been achieved.However,most of the current spatio-temporal miningrelated work is still in a coarse-grained spatio-temporal setting,which yields relatively limited social value and cannot be directly applied to fine-grained spatio-temporal tasks.There is still much room for improvement regarding fine-grained spatio-temporal data mining algorithms.(1)urban computing is a comprehensive problem,and it is necessary to introduce cross-domain data to assist in the analysis.However,existing cross-domain data fusion methods are post-fusion and ignore the joint features of cross-domain data at similar time slots,which directly leads to poor prediction performance;(2)fine-grained spatio-temporal settings exacerbate the problem of imbalance in the originally sparse classification labels,and such extremely sparse non-zero labels often render the neural network ineffective;(3)fine-grained spatio-temporal settings exacerbate the missing data(3)the fine-grained spatio-temporal setting will aggravate the problem of missing data,and the dataset with missing values will also greatly affect the algorithm’s extraction of features from the data distribution,making the final prediction results inaccurate.To address these challenges,this paper investigates the key techniques of data mining algorithms for fine-grained spatio-temporal settings in urban computing,using the powerful high-order feature extraction capability of deep learning models and graph neural network algorithms that fit well with the urban topology space to mine these spatio-temporal big data.This paper will focus on "how to effectively fuse cross-domain data to support fine-grained spatio-temporal data prediction","how to analyse and predict zero-inflated data exacerbated by fine-grained spatio-temporal settings" and "how to accurately impute missing data exacerbated by fine-grained spatio-temporal settings".The main contents and contributions of this paper are as follows.1.To address the problem of how to effectively fuse cross-domain data to support fine-grained spatio-temporal data prediction,this paper proposes a deep learning-based spatio-temporal cross-domain fusion algorithm to fuse cross-domain data to support the fine-grained spatio-temporal data prediction problem.The algorithmic model can effectively collect the spatio-temporal features of cross-domain data and the joint features between cross-domain data.Unlike previous post-fusion-based algorithms,the cross-domain fusion algorithm proposed in this paper can focus on the joint features of cross-domain data under similar time gaps through the cross-domain convolutional layer in the early stage of the model.This paper validates the algorithm with a typical urban anomaly detection scenario in urban computing that requires multi-domain feature fusion techniques and compares it with existing cross-domain fusion algorithms.The experimental results show the effectiveness of the method,which can achieve higher accuracy and recall in anomaly detection problems.2.To address the problem of how to analyse and predict sparse data further exacerbated by fine-grained spatio-temporal settings,this paper proposes a sparse label enhancement algorithm based on discrete data continuum,and adds an attention mechanism to the above cross-domain fusion model so that the model can complete the task of analysing and predicting spatio-temporal data under extremely sparse data conditions.The proposed continuousisation algorithm for discrete data converts the original discrete data into continuous data by performing linear transformation and logarithmic operations on the data,while injecting non-zero element information from the history of the spatial nodes corresponding to the labelled elements into the continuous data.The attention mechanism also allows the model to better extract the spatio-temporal characteristics of the cross-domain data.The algorithm is validated on an urban crime prediction scenario with extremely sparse events.The experimental results show that this discrete data continuum algorithm can significantly improve the performance of the neural network model in the presence of categorical label imbalance,and in combination with the fusion model based on the attention mechanism can accomplish spatio-temporal data analysis and prediction under sparse data labeling conditions.3.To address the problem of accurate data completion for missing data exacerbated by fine-grained spatio-temporal settings,this paper proposes a data completion algorithm based on a combination of graph attention mechanism and meta-learning mechanism.The algorithm complements the missing data by mining the spatio-temporal correlation of the data.The model first extracts the potential spatial relationships between nodes by combining graph embedding techniques with meta-learning.A priori knowledge is added to the auxiliary tasks to constrain and correct the results of the complementary data.This paper compares existing data completion algorithms with missing traffic data scenarios in smart transportation applications.Real data sets collected in two cities,Shanghai and Shenzhen,as well as simulated data sets based on various scenarios are used in the experiments.The experimental results show that the proposed algorithm gives more accurate results for data completion and shows sufficient robustness in different scenarios.
【Key words】 Urban computing; Spatiotemporal data mining; Deep learning; Graph neural network;
- 【网络出版投稿人】 北京邮电大学 【网络出版年期】2024年 01期
- 【分类号】TP311.13