节点文献

复杂函数型数据的有效统计分析

Effective Statistical Analysis of Complex Functional Data

【作者】 刘华;

【导师】 尤进红; 曹际国;

【作者基本信息】 上海财经大学 , 数理统计学, 2021, 博士

【摘要】 当一个变量被多次测量或观察时,这个变量可以看作是一个函数,则该变量称为函数型变量,而该变量的数据称为函数型数据。函数型数据通常是时间的函数,但也可能是空间位置,波长等的函数。函数型数据作为统计学的一个新领域,近些年是很多国内外学者关注的热点并且在很多领域得到了广泛的应用,例如临床,生物统计学,流行病学,社会和经济领域。在函数型数据分析中,函数型回归模型刻画了函数型变量和标量变量之间的关系。根据响应变量和协变量的类型,通常可以将函数型回归模型分为三类:(i)响应变量是函数,协变量是函数(function-on-function);(ii)响应变量是标量,协变量是函数(scalar-on-function);(iii)响应变量是函数,协变量是标量(function-on-scalar)。随着研究问题复杂性的增加,函数型数据出现了一些新的特征,包括大规模性、动态性、交互性等。为此需要提出新的模型和分析方法。本文基于不同的函数型回归模型研究了几个问题。首先,受近期研究大量函数型数据(例如COVID-19数据)的启发,本文提出了一种新的动态交互半参数函数型回归模型。该模型研究了一组协变量之间的动态交互效应及其对函数型响应变量的影响。该模型包括了最近提出的许多重要模型。通过张量积B-样条近似未知的二元系数函数,本文提出了一个三步估计法来迭代地估计未知的两元变系数函数,单指标参数的向量以及随机函数的协方差函数。本文还对得到的估计的渐近性质进行了研究,包括收敛速度和渐近正态分布。此外,本文还基于L2距离构造了一个检验统计量来检验动态交互效应是否随时间/空间位置而变化,并证明了该检验统计量的渐近正态性。我们通过三个数值模拟研究了本文提出的方法的表现和假设检验的表现。本文还将提出的动态交互半参数函数型回归模型用于分析COVID-19数据和ADNI数据。在这两个实际应用中,假设检验的结果表明,两元变系数函数随单指标和时间/空间位置而显着变化。例如,我们发现人口老龄化与社会经济协变量(例如每1000人中的病床数,医生,护士和人均GDP)的交互效应对COVID-19的死亡率的影响在COVID-19大流行的不同时期有所不同。还通过用动态交互半参数函数型回归模型估计了 141个国家/地区与COVID-19死亡率相关的医疗设施水平指数。其次,随着科学技术的发展,数据量呈指数增长,这为研究人员提供了更多的分析信息。同时,尽管计算资源迅速发展,但是大量的数据也给研究人员分析数据带来了挑战。一个挑战是,使用海量数据拟合模型需要太多内存,甚至超过了一台计算机的最大容量。此外,计算时间太长,无法获得结果。为了解决这些难题,其中一个有效的方法是从海量数据中抽取子样本作为全部数据的代理替来进行分析。受函数型线性模型中海量数据的存储和计算的挑战的启发,本文通过最小化抽样估计量与基于完整数据得到的估计量的渐近积分均方误差(IMSE),为函数型线性模型提出了一种基于L-最优准则的最优抽样方法。与使用所有数据进行计算相比,该算法具有较高的计算效率,并大大减少了计算时间。此外,本文给出了抽样估计量的渐近性质。在数值模拟中,分别在三种情况下,对本文提出的抽样方法的表现以及其与均匀抽样方法的比较进行了研究。另外,我们使用此抽样算法对三个阶段的全球气候数据进行了分析,每个阶段的数据的样本量均为n=1,028,032。通过对全球气候数据的分析,很明显可以看出基于L-最优准则得到的最优抽样方法比均匀抽样方法表现要好,并且可以很好地对基于完整数据得到的估计结果进行近似。第三,对于函数型广义线性模型,本文也提出了一种基于L-最优性准则的最优抽样方法来解决这些计算时间和存储难题。本文还对函数型广义线性模型下的通过抽样方法获得的估计量的渐近性质进行了研究。在数值模拟中,本文分别在函数型逻辑回归和函数型泊松回归两种情况设定下,对本文提出的函数型广义线性模型下的最优抽样方法的表现进行了研究,并将其与均匀抽样方法进行了比较。此外,本文使用函数型广义线性模型下的抽样方法来分析肾脏移植数据。数据模拟以及肾脏移植数据的结果都表明,本文提出的最优抽样方法要优于均匀抽样方法,并且最优抽样的结果非常接近基于完整数据得到的估计结果。最后,我们把本文中提出的模型和分析方法所涉及的算法都编成了相应的R语言代码和软件包,以便其他研究者使用。

【Abstract】 When a variable is measured or observed at multiple times,the variable can be treated as a function of time.The variable is therefore called functional variable,and the data for the variable are called functional data.The functional data are often functions of time,but may also be functions of spatial locations,wavelengths,etc.In recent years,functional data analysis has received considerable attention in many applied fields such as in clinical,biometrical,epidemiological,social and economic fields.Functional regression models the relationship among functional and scalar variables,and is widely used in functional data analysis.In the existing literature about functional regression,we can divide them into three categories depending on whether the responses or covariates are functional or scalar data:(ⅰ)functional responses with functional covariates;(ⅱ)scalar responses with functional covariates;and(ⅲ)functional responses with scalar covariates.With the increase in the complexity of real data,some new characteristics of functional data have emerged,including large-scale,dynamic,interact effect,and so on.Thus,it is necessary to propose new models and methods.In this dissertation,we study some problems based on different functional regression models.Firstly,motivated by recent work studying massive functional data,such as the COVID-19 data,we propose a new dynamic interaction semiparametric function-on-scalar(DISeF)model.The proposed model is useful to explore the dynamic interaction among a set of covariates and their effects on the functional response.The proposed model includes many important models investigated recently as special cases.By tensor product B-spline approximating the unknown bivariate coefficient functions,a three-step efficient estimation procedure is developed to iteratively estimate bivariate varying-coefficient functions,the vector of index parameters,and the covariance functions of random effects.We also establish the asymptotic properties of the estimators including the convergence rate and their asymptotic distributions.In addition,we develop a test statistic to check whether the dynamic interaction varies with time/spatial locations,and prove the asymptotic normality of the test statistic.The finite sample performance of our proposed method and the test statistic is investigated with three simulation studies.Our proposed DISeF model is also used to analyzing the COVID-19 data and the ADNI data.In both applications,hypothesis testing shows that the bivariate varying-coefficient functions significantly vary with the index and the time/spatial locations.For instance,we find that the interaction effect of the population ageing and the socio-economic covariates,such as the number of hospital beds,physicians,nurses per 1,000 people and GDP per capita,on the COVID-19 death rate varies in different periods of the COVID-19 pandemic.The healthcare infrastructure index related to the COVID-19 mortality rate is also obtained for 141 countries estimated based on the proposed DISeF model.Secondly,the volume of data increases exponentially with the development of science and technology,which provides researchers more information to analyze.At the same time,despite the rapid development of computational resources,the extraordinary amount of data also brings some challenges to researchers in analyzing data.One challenge is that fitting a model using massive data needs too much memory to exceed the maximum capacity of a single computer.Moreover,the computing time is too long to obtain the results.To tackle these challenges,an effective way is to take random subsamples from the massive data as a surrogate.Motivated by the memory and computation challenges in the massive data in the scalar-on-function linear model,we propose an optimal subsampling method based on L-optimality for functional linear model through minimizes the asymptotic integrated mean squared error(IMSE)of subsampling estimator in approximating estimator based on full data.This algorithm is computationally efficient and has a significant reduction in computing time compared to the full data approach and we establish the asymptotic properties of the subsampling estimators.The finite sample performance of our proposed method and the comparisons with the uniformly subsampling method are investigated simulation studies under three scenarios.Also,we use this algorithm to analyze the global climate data from three stages with full data size n=1,028,032.The results from the analysis of this data set show that the optimal subsampling method motivated by the L-optimality criterion is better than the uniform subsampling method and can well approximate the results based on full data.Thirdly,for the scalar-on-function generalized linear model,we also propose an optimal subsampling method based on the L-optimality criterion to tackle these computing time and memory challenges.We also establish the asymptotic properties of the estimators obtained by the subsampling method.The finite sample performance of our proposed subsampling method is investigated with two simulation studies under functional logistic regression and functional Possion regression,respectively.And,we use the kidney transplant data to illustrate our proposed subsampling method for the functional generalized linear model.The finite sample performance and results of the empirical application with kidney transplant data show that our subsampling denotes the uniformly subsampling method and can well approximate the results based on full data.R code and R package have been developed for implementing the proposed methods.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络