节点文献
高维潜因子模型的理论和应用研究
Theory and Application of High-dimensional Latent Factor Models
【作者】 刘伟;
【导师】 林华珍;
【作者基本信息】 西南财经大学 , 统计学, 2020, 博士
【摘要】 自本世纪初以来,科学技术的进步使高维数据在遗传学、分子生物学、认知科学、环境科学、天体物理学、金融和互联网商务等领域呈现爆炸性的增长。相关性是高维数据的本质特征之一,而因子模型就是基于变量之间的相关性来进行主要信息提取的常用方法,已被证实在构建观测数据的同质性和相关性方面卓有成效(Fan et al.,2017),因而成为数据压缩和降维的重要工具。然而,现有的高维因子模型主要有以下不足:首先,因子模型解释性较差。传统方法中,每个载荷的非零估计导致因子与所有高维变量都有关,无法解释每个因子由哪些变量贡献产生,因子的意义不清楚。比如在DNA甲基化水平数据中,我们关心哪些Cp G位点对提取的因子有贡献,从而对有贡献的Cp G位点进行进一步的基因集富集分析。再次,现有的高维因子模型无法处理非连续型变量,比如二元变量和计数变量等。实际上,典型的高维数据,例如全基因组关联分析的基因数据集中,每个SNP都是取值为{0,1,2}的分类变量,传统的因子模型处理这样的数据是直接当做连续变量处理,忽略了变量类型携带的信息,可能导致错误的结果。最后,目前的因子模型只是作为数据压缩和降维的重要工具,用于无监督学习。如何将因子模型应用于有监督的回归模型,从而构建超高维强相关协变量对感兴趣的结果的影响,也是重要的研究课题。针对这些现存问题,本文构建了一系列模型和方法来逐一解决。为解决解释性问题,本文提出了高维稀疏因子模型及两步的判罚最小二乘估计方法。由于载荷的稀疏性、因子的不可观测性以及大样本容量和高变量维度,同时估计因子和载荷计算成本较高。为了降低计算复杂度,我们创造性地提出了一种非迭代的两步方法,其中每一步都有显式解。所得载荷矩阵的稀疏结构,能使我们了解每个因子由哪些变量贡献产生,从而解释每个因子的意义。一个自然的问题是得到的载荷矩阵的稀疏结构是正确的吗?为了进一步确定所得稀疏载荷矩阵的正确性,我们基于样本划分的思想提出了双向联合检验方法,从两个方向对载荷矩阵进行联合推断,一是检验其中的零载荷向量是否真的为零,二是检验其中的非零载荷向量中是否还存在零。基于双向联合检验,我们可以推断出与估计的潜因子相关的变量,以及与潜因子无关的变量,从而明确了各潜变量的含义,实现了因子的可解释性。理论上,我们证明了载荷估计具有Oracle性质,并且证明联合检验方法渐近地达到预先指定的显著性水平,功效渐近地逼近1。仿真研究和实际数据分析,验证了我们方法的有效性。随着技术的进步,收集的高维数据越来越复杂。特别是高维数据不再以连续这种单一类型呈现,混合高维数据逐渐普及和常见。使用一组低维因子来表示混合类型的高维数据具有重要的理论及应用价值。对于现有高维因子模型只能处理连续变量的局限性,本文针对超高维混合类型变量,提出了一种广义因子模型(GFM)及其相应的算法和理论。特别,为了解决非线性和混合类型变量带来的计算问题,我们设计了一个两步方法,使得每步更新都可以使用现有的软件包沿着变量和样本两个方向来并行计算。从理论上确立了了当样本量n和变量维数p都发散到无穷大且非线性结构随混合变量类型而变化时,因子和载荷估计的收敛速度。此外,因子数目的正确选择对于因子模型的理论有效性和经验有效性都至关重要,我们因此还提出了一个基于判罚损失的准则来估计广义因子模型框架下的因子个数,并从理论上证明该估计是相合的。广泛的模拟研究证实我们的方法相对于现有方法有显著的优势。基于对基因数据和心律失常数据的分析结果,我们得到了比现有因子模型更具预测性和解释性的载荷和因子估计。纯粹的因子模型只能用于无监督学习,大大限制了因子模型的应用范围。考虑到因子模型处理超高维强相关变量的天然优势,本文还探索了因子模型和回归模型的结合,来解决传统的高维回归模型无法处理的超高维强相关协变量的问题,并提出了一种基于潜因子的超高维相关协变量半参数多指标模型。目前处理超高维变量的回归模型大致有两种方法,一是仅考虑响应变量与协变量关系,如筛选和变量选择方法,这种方法不能处理变量之间有强相关的情形;二是只考虑协变量信息,如主成分方法提前特征,压缩数据。但提取的特征可能与感兴趣的结果变量无关。我们提出的方法基于响应变量与协变量之间的关系,从大量的预测变量中提取低维潜在特征,因此提取的特征是有监督的,可以表达响应变量与协变量之间的关系。此外,我们不需要设定响应变量的分布,允许响应变量和潜在因子之间的连接函数未知,因此有很强的数据适应性。进一步,我们提供了一个基于似然函数的框架来估计参数,并且证明所得估计是相合、渐近正态及有效。最后,由于高维参数的估计在每一步都有显式形式,所提的方法计算简单且有效。模拟研究表明,我们提出的方法在灵活性和有效性方面都优于其他现存方法。将该方法进一步应用于基因表达数量性状位点研究,与现有方法相比,得到了额外的18个SNPs与肺组织中目标基因表达有关。
【Abstract】 Since the beginning of this century,the progress of science and technology has led to the explosive growth of high-dimensional data in genetics,molecular biology,cognitive science,environmental science,astrophysics,finance,Internet commerce and other fields.Correlation is one of the essential features of high-dimensional data,and the common-used method for extracting the main information based on the correlation between variables is the factor model,which has been proved to be effective in simultaneously modeling the commonality and cross-sectional dependence of the observed data(Fan et al.,2017),so it has once again become a powerful framework for data compression and dimensionality reduction.However,the existing high-dimensional factor models mainly have the following shortcomings.Firstly,the interpretability of the factor model is poor.In traditional methods,the non-zero loading estimate leads to factors related to all high-dimensional variables,so it is impossible to explain which variables contribute to each factor,and the meaning of the factor is not clear.For example,in the DNA methylation level data,we are concerned about which Cp G sites contribute to the extracted factors,so as to carry out further gene set enrichment analysis for the contributed Cp G sites.Secondly,the existing high-dimensional factor model can not deal with non-continuous variables,such as binary variables and count variables.For example,in the gene dataset of genome-wide association analysis,each SNP is a categorical variable takeing values of {0,1,2}.The traditional factor model treats such data directly as continuous variables,ignoring the information carried by the variable type,which may lead to wrong results.Finally,the current factor model can only be used for unsupervised learning as a tool for data compression and dimensionality reduction.It is also an important research topic to apply the factor model to the supervised regression model to explore the effect of ultra-high-dimensional strong correlation(UDSC)covariates on the outcome of interest.In view of these existing problems,we establish appropriate models and puts forward appropriate methods to solve them one by one.To solve the interpretation problem,a high-dimensional sparse factor model is established and a two-step penalized least squares estimation method is developed.Because of the sparsity of loading,the unobservability of factor and the large sample size and high variable dimension,the cost of computation is very high to simultaneously estimate factor and loading.To reduce the computational complexity,we creatively propose a non-iterative two-step method,in which each step has an explicit solution.The sparse structure of the resulting loading matrix enables us to understand which variables contribute to each factor,thus explaining the meaning of each factor.A natural question is whether the obtained sparse structure of the loading matrix is correct.To confirm the resulting loadings estimator,a simultaneous testing procedure is designed to make the simultaneous inference on a set of loading vectors from two directions based on the sample-split idea,where one is testing whether selected zero loading vectors indeed are zero and the other is testing whether there exists zero in the given nonzero loading vectors.As a result,we can fully determine the variables which are associated with an estimated latent factor,as well as the variables which are independent of the latent factor.Consequently,the meaning of each latent variable is clearly figured out and the interpretability of factors is achieved.In theory,we prove that the loadings estimator has the oracle property,and the simultaneous testing procedure asymptotically achieves the pre-specified significance level and has the testing power converging to one.The performance and effectiveness of our method are demonstrated via simulation studies and real data analysis.With the advancement of technology,the collected high-dimensional data becomes more and more complex.In particular,high-dimensional data is no longer presented in a single continuous type,and mixed high-dimensional data is gradually becoming popular and common.Using a set of low-dimensional factors to represent mixed types of high-dimensional data has important theoretical and application value.Due to the limitation of the existing methods for factor analysis that deal with only continuous variables,in this paper,we develop a generalized factor model,a corresponding algorithm and theory for ultra-high dimensional mixed types of variables.Specifically,to solve the computational problem arising from the non-linearity and mixed types,we develop a two-step algorithm so that each update can be carried out in parallel across variables and samples by using an existing package.Theoretically,we establish the rate of convergence for the estimators of factors and loadings in the presence of nonlinear structure accompanied with mixed-type variables when both sample size n and variable dimension p diverge to infinity.Moreover,since the correct specification of the number of factors is crucial to both the theoretical and the empirical validity of factor models,we also develop a criterion based on a penalized loss to consistently estimate the number of factors under the framework of a generalized factor model.Extensive simulation studies have confirmed that our method has significant advantages over existing methods.Based on the analysis results of genetic data and arrhythmia data,we have obtained more predictive and interpretable estimators for loadings and factors than the existing factor models.The pure factor model can only be used for unsupervised learning,which greatly limits the scope of application of the factor model.Considering the natural advantage of factor model in dealing with UDSC variables,this paper also explores the combination of factor model and regression model to solve the regression problem with UDSC covariates that can not be processed by traditional high-dimensional regression models,and a semi-parametric multi-index model of UDSC covariates based on latent factors is proposed.At present,there are roughly two methods to deal with the regression model of ultra-highdimensional variables.One is to only consider the relationship between response variables and covariates,such as screening and variable selection methods.This method cannot handle situations where there is a strong correlation between variables;The other is to only consider the covariates information,such as principal component method to extract features and compress data.But the extracted features may not be related to the outcome variable of interest.Our proposed method is based on the relationship between response variables and covariates,and extracts low-dimensional potential features from the vast predictive variables.Therefore,the extracted features are supervised and can express the relationship between response variables and covariates.In addition,the proposed method is very flexible,since we allow the distributions of both the covariates and response,as well as the forecasting function between the target and the latent factors to be unknown.Furthermore,we provide a framework based on the likelihood function to estimate the parameters,and prove that the obtained estimator are consistent,asymptotically normal and efficient.Finally,since the estimates for the high dimensional parameters have a closed form in each step,the computation and programming are simple.Extensive simulation studies demonstrate that our proposed method outperforms other competitors in terms of applicability and empirical efficiency.The application of the proposed method is further applied to analyze the e QTL study,resulting in 18 SNPs that are not detected by existing methods.
【Key words】 High-dimensional; Strong correlation; Latent factor; Nonlinear; Estimation; Simultaneous inference;