节点文献

矩阵变量下的双向因子模型

Matrix-variate Data Analysis by Two-way Factor Models

【作者】 李岩;

【导师】 郭建华;

【作者基本信息】 东北师范大学 , 机器学习与生物信息学, 2023, 博士

【摘要】 随着大数据时代的来临,我们获取数据的途径越来越丰富,可获得数据的形式也更加多样化。从观测的随机数,到多元的向量数据,再到多个维度的矩阵数据或张量数据。从随机变量,到随机向量,再到随机矩阵。关于随机矩阵的研究逐渐成为统计学关注的重点。例如经济学中对面板数据的研究,网络数据中的社区探测,生物信息学中对基因阵列的探测,图像和信号的处理等等。然而矩阵变量之间的相关性要远比向量型数据更加复杂,这就导致我们不能简简单单地用随机向量的研究方法来研究矩阵变量数据。因此发展基于矩阵的统计方法是很有必要的。因子模型是多元统计中一个重要的模型,它最早提出就是为了分析多个变量之间的相关关系。随着变量维数的增加,因子分析模型被广泛的应用到了数据降维以及对变量进行解释等方面,并且取得了很丰富的成果。我们基于向量型因子模型,并将单一观测矩阵数据的双向因子模型扩展到含有多次重复观测矩阵变量数据的分析研究中。为了区分之前单一矩阵的双向因子模型(2wFM),将我们的模型称之为矩阵变量的双向因子模型(2wFMs)。和2wFM类似,我们采用了一种可加的因子模型结构分别刻画矩阵行变量之间以及列变量之间的相关性,我们假定这种行相关性和列相关性是可以分开的。具体地,由于因子模型是对相关阵进行建模,我们的模型将总体的协方差阵分解成两个低秩矩阵和一个对角矩阵的加和,这三部分分别表示行变量、列变量以及特殊因子的方差。基于2wFMs,我们讨论了参数的可识别性,并在可识别性的条件下,给出了参数的极大似然估计。在固定维数的情况下,我们同时给出了基于主成分的估计方法,并把它作为我们算法的初值选取,这大大降低了我们算法所需要的时间。理论方面,我们分别针对固定维数情况(变量维数p,q固定,样本量n趋于无穷)和高维情况(n,p,q趋于无穷),对极大似然估计的相合性和渐近正态性给出了理论上严格的证明。虽然模型条件在固定维数和高维情况下是相同的,但理论上完全是两个不同的证明方向。对于固定维数,我们采用了基于经典的M-估计和Δ-方法来证明参数的相合性和渐近正态性;而对于高维情况,渐近正态性的证明则是基于估计方程之间的关系。这也为高维统计理论分析和固定维数的传统统计推断的区分提供了一个思路。针对变量维数固定和变量维数趋于无穷两种情况,我们分别给出了一系列的数值模拟来验证估计精度随观测样本量以及变量维数变化时的变化情况,这些模拟的结果和我们的理论结果是相吻合的。另外,将我们的模型应用到了经济合作与发展组织(OECD)的关键经济指标(KEI)的数据以及进出口贸易数据(Direction of trade statistics(DOTS))上,并且通过我们的模型给出变量关系的一些解释。

【Abstract】 In the era of big data,we can now collect data in a variety of ways,and more and more types of data are becoming available.From random variables to random vectors,even to random matrices.Such as consideration of development trends in economic and financial data,community recognition in network data,detection in gene microarray,picture and signal processing,and so on.Matrix observations are becoming increasingly important to statisticians.However,because the correlation among matrix variables is much more complicated than that of vector-variate data,vectorized observations will lose a lot of information about the matrix’s internal structure.As a result,a statistical method based on the matrix is necessary.The factor model is an essential model in multivariate statistics.It was first proposed for the purpose of investigating the relationship between variables.The factor model has been widely used in dimensionality reduction,and a number of studies have relied on it.Based on the factor model,we extended the bidirectional factor model of a single observation to the analysis of matrix-variate data with duplicated observations.To distinguish our model from the earlier 2-way factor model of a single matrix(2wFM),we named our model with duplicate observations as 2wFMs.Similarly to 2wFM,we use an additive factor model structure to describe the correlation between matrix variables’ rows and columns,assuming that the row and column correlations can be separated.Our model decomposes the covariance matrix into the sum of two lowrank matrices and a diagonal matrix,which represent the correlation among the row variables,column variables,and the variance of idiosyncratic components,respectively.We explore the identifiability of parameters through 2wFMs and present the maximum likelihood estimate of parameters under the condition of identifiability.We also provide an estimating approach based on principal components in the case of fixed dimensions,which is chosen as the initial value of our algorithm,considerably reducing the time it takes.In terms of theory,we’re focusing on the fixed-dimension(p,q are fixed,n goes to infinity)and high-dimensional case(n,p,q all go to infinity),respectively.Even if the model conditions in fixed and high dimensions are similar,they are theoretically two different proving methods.To verify the parameter consistency and asymptotic normality in fixed dimensions,we apply the traditional M-estimation and Δ-method,whereas in high dimensions,the proof of asymptotic normality is based on the relationship between the estimation equations.As a result,it gives a way to identify high-dimensional statistical theoretical analysis from typical fixed-dimensional statistical inference.We provide numerical simulations for these two situations to evaluate the precision of estimators under different sample sizes and variable dimensions.The results are consistent with our theoretical results.In addition,we applied our model to data from the Key Economic Indicators(KEI)data of the Organization for Economic Cooperation and Development(OECD)and the Direction of Trade Statistics(DOTS)dataset corresponding to the interpretation of the variables.

  • 【分类号】O151.21;O212
节点文献中: