节点文献

非线性时间序列的复杂性和相似性研究及其应用

Complexity and Similarity Studies of Nonlinear Time Series and Their Applications

【作者】 王卓;

【导师】 商朋见;

【作者基本信息】 北京交通大学 , 统计学, 2025, 博士

【摘要】 在当今科学技术飞速发展的时代,复杂非线性时间序列普遍存在于金融、工业、医疗等各个领域.这些序列蕴含着丰富的信息,有效提取序列所包含的信息并分析特征,对于理解系统的内在运行机制、预测未来行为、模式识别以及故障检测具有重要意义.尽管针对非线性时间序列的研究已经取得一系列成果,但面对复杂多样的数据特点、多样化的问题场景,以及对分析精度和效率的更高要求,非线性时间序列的研究仍然面临诸多挑战.此外,在社会持续发展进程中,不断推进理论和方法的创新,为解决各领域的实际问题提供更有效的技术方案至关重要.本文重点针对非线性时间序列的复杂性和相似性分析中面临:如何有效降低噪声和模型参数等因素对算法性能的干扰、如何更全面地描述系统的结构特性、如何准确地捕捉高维序列的动态行为以及多变量的协同变化、如何有效处理具有非线性流形结构的复杂数据等问题展开深入研究.我们从信息熵、复杂性熵平面、能量距离、黎曼几何多视角出发,并结合符号动力学分析方法、相空间重构理论、复杂网络和机器学习算法构建数学模型,旨在能够更高效地度量非线性时间序列的复杂性特征,提高序列间相似性分析的准确性和效率.本文的主要研究内容和贡献可归纳为以下四个方面:1.基于累积剩余信息熵的时间序列复杂性分析.首先,我们提出累积剩余Tsallis奇异熵(Cumulative Residual Tsallis Singular Entropy,CRTSE)模型,旨在提高非线性序列复杂性分析的有效性与鲁棒性.CRTSE模型运用奇异值来表征序列的概率分布,突破累积剩余熵对变量非负的限制,有效降低噪声干扰.同时,CRTSE保留了Tsallis熵的非广延性,可通过灵活调整参数精准地刻画复杂系统结构.通过模拟实验对比CRTSE与Tsallis奇异熵的性质,证实CRTSE能够更高效地提取复杂系统的动态特征.为进一步挖掘CRTSE在分类和预测中的潜力,我们构建基于CRTSE的灰狼优化支持向量机(CRTSE-GWOSVM).该方法先运用CRTSE提取信息特征,再借助灰狼优化器(Grey Wolf Optimizer,GWO)优化SVM参数,进而实现复杂系统的精准分类.在铁路轨道故障和工业轴承故障的检测识别任务中,CRTSE-GWOSVM相较于基于网格搜索优化、粒子群优化的SVM,展现出更为出色的分类性能.本研究成果在复杂系统智能分类领域具有广阔的应用前景.2.基于离散模式的复杂性-累积剩余Tsallis熵平面分析.离散熵是近年来提出的一种快速且准确的符号动力学分析方法.我们引入离散熵计算序列的概率分布,创新性地提出基于离散熵的复杂性-累积剩余Tsallis熵平面(Complexity-Cumulative Residual Tsallis Entropy Plane,C-CRTEP),能够比单一的熵指标更全面综合地剖析复杂系统内在结构.模拟实验结果表明,基于离散熵的C-CRTEP能够准确刻画数据结构特征.通过绘制随非广延参数变化的曲线,能显著区分随机过程与混沌系统.进一步地,我们将基于离散熵的C-CRTEP算法拓展至高维空间,提出基于多元离散熵的C-CRTEP.我们通过实验证实了该模型在多元序列特征提取方面的有效性.在金融股票市场的特征分析与识别区分中,基于离散模式的C-CRTEP均表现出优异的性能.3.基于能量距离理论的时间序列复杂性分析.信息熵计算通常依赖于序列的概率分布.然而,这一过程不仅可能会导致信息丢失,还对数据长度有一定要求.此外,如何准确度量高维序列的复杂性特征以及多变量的协同变化,一直是该领域研究的重点与难点.针对上述问题,我们聚焦于能量距离理论来探究非线性系统的复杂性特征.距离分量作为基于能量距离的多样本能量统计量,在高维空间具备旋转不变性,处理复杂数据优势显著.我们将距离分量与相空间重构理论结合,提出广义距离分量(Generalized Distance Components,GDISCO)算法,通过衡量序列离散程度来反映其复杂性行为.该算法的优势及创新之处在于,GDISCO对一维和高维序列采用统一的定义方式.对于高维序列,GDISCO不仅给出了总离散度统计量,还涵盖组内和组间离散度统计量.经一维和高维数据模拟实验验证,GDISCO能准确刻画并区分不同数据的结构复杂性,以及多变量间的协同关系.为进一步提升算法性能,我们引入权重信息对GDISCO进行拓展.为此,我们构建了一种新的空间振幅和趋势差权重算子(Spatial Amplitude and Trend Difference Weighting Operator,WSAT D),其可通过调整参数β为数据分析提供更多的灵活性与适用性.然后,我们提出基于WSAT D的GDISCO方法,模拟实验结果表明,WSAT D-GDISCO在描述系统的结构特性方面具有更大的潜力.在生理信号和金融股票市场的研究中,GDISCO及其加权算法均取得了优异的表现.本研究成果为高维复杂系统的分析提供了新视角与方法.4.基于黎曼几何的时间序列不可逆性和相似性研究.针对具有非线性流形结构的复杂数据,经典欧氏度量难以精准描述数据点之间的关系.为此,我们从黎曼几何角度剖析时间序列的复杂性特征,并探究序列间的相似性关系.一方面,我们创新性地提出基于仿射不变黎曼度量的时间不可逆性度量(Time Irreversibility Measure Based on Affine Invariant Riemannian Metric,TIAIRM),用于刻画复杂系统的动态行为.TIAIRM借助滑动窗口和相空间重构理论,将时间序列映射至黎曼流形进行分析,能精确获取序列的非线性流形结构特征,并给予几何层面解释.该方法通过滑动窗口使子序列满足弱平稳条件,对非平稳序列同样适用,且可度量局部及全局时间不可逆性.经模拟数据与替代数据实验证实了TIAIRM的有效性.在金融股票市场的分析中,TIAIRM成功识别出股票市场的金融危机时期和稳定时期.在生理信号的应用中,TIAIRM准确检测出不同类型心电信号、不同患病程度的帕金森患者步态数据的时间不可逆性缺失程度.另一方面,我们创新性地融合复杂网络与黎曼几何理论,提出基于序数网络的仿射不变黎曼度量(Ordinal Network-based Affine Invariant Riemannian Measure,ONAIRM),用于量化时间序列的相似性与差异性.利用图论思想将时间序列转化为复杂网络,在刻画系统动态行为方面具有显著优势.序数网络通过构建关于序数模式的有向网络,可充分挖掘序列潜在的复杂动态模式.模拟数据验证了ONAIRM的有效性和对噪声的鲁棒性.在此基础上,我们进一步提出基于ONAIRM的古典多维标度算法(ONAIRM-CMDS),该算法能度量序列差异性进行可视化呈现.相较于传统方法,ONAIRM-CMDS能更精准区分周期序列、随机过程与混沌系统.此外,为了更好地开展分类和预测任务,我们还构建了基于ONAIRM的层次聚类算法和基于ONAIRM的黎曼均值分类算法.将上述拓展算法应用于生理信号分类和铁路轨道故障检测的研究中,均展现出优异的性能.研究结果表明,ONAIRM在复杂系统分类与预测领域具有重要的应用价值.

【Abstract】 In today’s era of rapid scientific and technological advancements,complex nonlinear time series are widely present in various fields such as finance,industry,and healthcare.These time series contain abundant information,and effectively extracting and analyz-ing their features is crucial for understanding the intrinsic operational mechanisms of systems,predicting future behaviors,pattern recognition,and fault detection.Although significant progress has been made in the study of nonlinear time series,numerous chal-lenges remain due to the complexity and diversity of data characteristics,the wide range of problem scenarios,and the increasing demands for analysis accuracy and efficiency.Furthermore,as society continues to evolve,it is essential to continuously advance theo-ries and methodologies to provide more effective technological solutions for real-world problems in various domains.This paper focuses on the challenges in the complexity and similarity analysis of nonlinear time series,including how to effectively reduce the interference of noise and model parameters on algorithm performance,how to comprehensively describe the struc-tural characteristics of the system,how to accurately capture the dynamic behavior of high-dimensional sequences and the collaborative variations of multiple variables,and how to efficiently handle complex data with nonlinear manifold structures.Starting from the perspectives of information entropy,complexity-entropy plane,energy distance,and Riemannian geometry,and integrating symbolic dynamics analysis methods,phase space reconstruction theory,complex networks,and machine learning algorithms,we construct a mathematical model aimed at more efficiently measuring the complexity characteristics of nonlinear time series and improving the accuracy and efficiency of similarity analysis between sequences.The main research contents and contributions of this paper can be summarized in the following four aspects:1.Complexity analysis of time series based on cumulative residual information entropy.Firstly,we propose the cumulative residual Tsallis singular entropy(CRTSE)model to enhance the effectiveness and robustness of nonlinear time series complexity analysis.The CRTSE model employs singular values to characterize the probability dis-tribution of a sequence,overcoming the constraint of non-negativity in cumulative resid-ual entropy and effectively reducing noise interference.Meanwhile,CRTSE retains the non-extensive property of Tsallis entropy,allowing for precise characterization of com-plex system structures through flexible parameter adjustments.By conducting compar-ative simulations between CRTSE and Tsallis singular entropy,we confirm that CRTSE can more efficiently extract dynamic features of complex systems.To further explore CRTSE’s potential in classification and prediction,we develop a grey wolf optimized support vector machine based on CRTSE(CRTSE-GWOSVM).This method first uti-lizes CRTSE to extract informative features,then applies the grey wolf optimizer(GWO)to optimize SVM parameters,thereby achieving accurate classification of complex sys-tems.In fault detection and identification tasks for railway tracks and industrial bearings,CRTSE-GWOSVM demonstrates superior classification performance compared to SVMs optimized via grid search and particle swarm optimization.The findings of this study hold promising application prospects in the field of intelligent classification for complex systems.2.Analysis of the complexity-cumulative residual Tsallis entropy plane based on dispersion patterns.Dispersion entropy is a recently proposed symbolic dynamics anal-ysis method that is both fast and accurate.We introduce dispersion entropy to com-pute the probability distribution of a sequence and innovatively propose the complexity-cumulative residual Tsallis entropy plane(C-CRTEP)based on dispersion entropy.Com-pared to single entropy metrics,C-CRTEP provides a more comprehensive analysis of the intrinsic structure of complex systems.Simulation results demonstrate that C-CRTEP based on dispersion entropy can accurately characterize data structural features.By plotting curves as a function of the non-extensive parameter,we can effectively distin-guish between stochastic processes and chaotic systems.Furthermore,we extend the C-CRTEP algorithm based on dispersion entropy to high-dimensional space,proposing the C-CRTEP based on multivariate dispersion entropy.Experimental results confirm the effectiveness of this model in extracting features from multivariate sequences.In the fea-ture analysis and classification of financial stock markets,C-CRTEP based on dispersion patterns consistently exhibits outstanding performance.3.Complexity analysis of time series based on energy distance theory.The calcula-tion of information entropy typically relies on the probability distribution of a sequence.However,this process may lead to information loss and imposes certain constraints on data length.Additionally,accurately measuring the complexity characteristics of high-dimensional sequences and their cooperative variations remains a key challenge in this field.To address these issues,we focus on energy distance theory to explore the com-plexity characteristics of nonlinear systems.Distance components,as multi-sample en-ergy statistics based on energy distance,possess rotational invariance in high-dimensional space,making them particularly advantageous for handling complex data.We integrate distance components with phase space reconstruction theory and propose the general-ized distance components(GDISCO)algorithm,which reflects the complexity behavior of a sequence by measuring its dispersion degree.The key advantage and innovation of GDISCO lie in its unified definition for both univariate and high-dimensional sequences.For high-dimensional sequences,GDISCO not only provides the total dispersion statistic but also includes the within-sample and between-sample dispersion statistics.Through simulation experiments on both univariate and high-dimensional datasets,GDISCO ef-fectively characterizes and differentiates structural complexity and the cooperative re-lationships among multiple variables.To further enhance algorithm performance,we introduce weighting information to extend GDISCO.Specifically,we develop a new spa-tial amplitude and trend difference weighting operator(WSAT D),which allows for greater flexibility and adaptability in data analysis by adjusting the parameterβ.Subsequently,we propose the WSAT D-GDISCO method.Simulation results indicate that WSAT D-GDISCO has greater potential in describing the structural characteristics of complex systems.In applications to physiological signals and financial stock market studies,both GDISCO and its weighted variant demonstrate outstanding performance.This research provides new perspectives and methodologies for analyzing high-dimensional complex systems.4.Study on irreversibility and similarity of time series based on Riemannian geom-etry.For complex data with nonlinear manifold structures,classical Euclidean metrics struggle to accurately describe relationships between data points.To address this chal-lenge,we analyze the complexity characteristics of time series and explore the similarity relationships between sequences from the perspective of Riemannian geometry.On one hand,we innovatively propose the time irreversibility measure based on affine invariant Riemannian metric(TIAIRM)to characterize the dynamic behavior of complex systems.TIAIRM utilizes sliding windows and phase space reconstruction theory to map time se-ries onto a Riemannian manifold for analysis,allowing it to precisely capture the nonlin-ear manifold structure of sequences and provide geometric interpretations.This method ensures that subsequences satisfy weak stationarity conditions through sliding windows,making it applicable to non-stationary sequences,and can measure both local and global time irreversibility.Experiments on simulated and surrogate data confirm TIAIRM’s ef-fectiveness.In financial stock market analysis,TIAIRM successfully identifies financial crises and stable periods.In physiological signal applications,TIAIRM accurately de-tects time irreversibility deficits in different types of electrocardiogram(ECG)signals and gait patterns of Parkinson’s patients with different disease severity.The ordinal network,by constructing a directed network based on ordinal patterns,can fully explore the underlying complex dynamic patterns of a sequence.On the other hand,we innovatively integrate complex networks with Riemannian ge-ometry and propose the the ordinal network-based affine invariant Riemannian measure(ONAIRM)to quantify the similarity and differences between time series.By leveraging graph theory,ONAIRM transforms time series into complex networks,which offer sig-nificant advantages in capturing system dynamics.Ordinal networks construct directed networks based on ordinal patterns,effectively uncovering hidden complex dynamic pat-terns within sequences.Experimental results on simulated data validate ONAIRM’s ef-fectiveness and robustness against noise.Building on this,we further propose the clas-sical multidimensional scaling algorithm based on ONAIRM(ONAIRM-CMDS),which measures sequence dissimilarity and provides a visual representation.Compared to tra-ditional methods,ONAIRM-CMDS more accurately distinguishes periodic sequences,stochastic processes,and chaotic systems.Additionally,to improve classification and prediction tasks,we develop the hierarchical clustering algorithm based on ONAIRM and a Riemannian mean classification algorithm based on ONAIRM.When applied to physi-ological signal classification and railway track fault detection,these extended algorithms demonstrate outstanding performance.The research results indicate that ONAIRM holds significant application value in the classification and prediction of complex systems.

  • 【分类号】O211.61
节点文献中: 

本文链接的文献网络图示:

本文的引文网络