节点文献

缺失数据情形三类统计模型的经验似然推断

Empirical Likelihood Inferences for Three Classes of Statistical Models with Missing Data

【作者】 黎玲

【导师】 秦永松;

【作者基本信息】 广西师范大学 , 概率论与数理统计, 2008, 硕士

【摘要】 缺失数据现象在现实生活中经常发生,如民意调查、市场调查、医药研究等领域常有数据缺失.在有数据缺失的情况下通常的统计方法往往不能直接应用,需要对数据进行必要的处理,处理带有缺失数据的不完全样本时常常需要对缺失值进行补充,继而得到“完全样本”,再按通常的统计方法进行推断,缺失数据情形的统计推断是当今统计界的一个热门研究领域(Little and Rubin, Statistical Analysis with Missing Data[M], New York: JohnWiley and Sons, 2002). Wang and Rao (Empirical likelihood for linear regression models underimputation for missing responses[J], Canadian J Statist, 2001, 29: 597-608)在固定设计及缺失数据情形构造了线性模型中回归系数的经验似然置信区间(域),他们采用的是通常的回归填补法补足缺失数据,再利用填补后的“完全样本”构造回归系数的经验似然比统计量,并证明了此经验似然比统计量的极限分布为加权卡方分布,在利用该结果构造回归系数的经验似然置信区间(域)时需要进行调整,从而需要估计调整系数,导致经验似然置信区间(域)精度的降低.我们在第二章中提出了一种新的数据补足方法,证明了基于此填补法得到的回归系数的经验似然比统计量的极限分布为卡方分布,利用此结果构造回归系数的经验似然置信区间(域)时不需要调整,从而可以提高经验似然置信区间(域)的覆盖精度.总体差异比较是医学、经济和教育领域经常遇到的课题,秦永松和赵林城(Semi-parametric likelihood confidence intervals for various differences of two populations[J], Statisticsand Probability Letters, 1997, 33(2): 135-143;两总体分位数差异的经验似然比置信区间[J],数学年刊, 1997, 18A(6): 687-694;两样本分位数差异的半经验似然比检验[J],应用数学学报, 1998, 21(1): 103-112;Empirical likelihood ratio confidence intervals for various differencesof two populations[J], System Science and Mathematical Sciences, 2000, 13: 23-30)在完全样本情形系统讨论了各种总体差异指标的经验似然置信区间的构造. Qin and Zhang(Empiricallikelihood confidence intervals for differences between two datasets with missing data[J], PatternRecognition Letters, 2008, 29(6): 803-812)在MCAR缺失机制下的不完全样本情形构造了两非参数总体差异指标的经验似然置信区间,他们采用的是单一随机填补法补足缺失数据,我们在第三章采用分数填补法补足缺失数据,在MCAR缺失机制下的不完全样本情形构造了两非参数总体差异指标的经验似然置信区间,由此可以提高置信区间的覆盖精度.在第四章中,将第三章的结果推广到MAR缺失机制情形,得到了MAR缺失机制下的不完全样本情形两非参数总体差异指标的经验似然置信区间.本文的特色体现在以下几个方面1.在讨论固定设计及缺失数据情形线性模型中回归系数的经验似然置信区间(域)的构造时,提出了一种新的数据补足方法,证明了基于此填补法得到的回归系数的经验似然比统计量的极限分布为卡方分布,利用此结果构造回归系数的经验似然置信区间(域)时不需要调整,从而可以提高经验似然置信区间(域)的覆盖精度.2.在MCAR缺失机制下的不完全样本情形,采用分数填补法(一种重复填补法)补足缺失数据,构造了两非参数总体差异指标的经验似然置信区间.通常的(单一)随机填补法是分数填补法的特例,当分数填补法中的重复次数增加时,可以逐步减少填补方差,与单一填补法比较,分数填补法可以提高置信区间的覆盖精度.3.在MAR缺失机制下的不完全样本情形,采用分数填补法(一种重复填补法)补足缺失数据,构造了两非参数总体差异指标的经验似然置信区间. MAR缺失机制比MCAR缺失机制的限制条件更弱且在实际中更易满足.

【Abstract】 Item non-response occurs frequently in daily life. It happens in opinion polls, market re-search surveys, medical studies and other scientific experiments. In such circumstances, the usualinferential procedures for complete data sets cannot be applied directly. It needs to do some treat-ments on data before we can use usual statistical approaches. A common method is to imputevalues for each missing response in order to obtain a‘complete sample’set and then apply stan-dard statistical methods. Statistical inference for missing data is an important research field (e.g.Little and Rubin, Statistical Analysis with Missing Data[M], New York: John Wiley and Sons2002). Wang and Rao (Empirical likelihood for linear regression models under imputation formissing responses[J],Canadian J Statist, 2001, 29: 597-608) obtain empirical likelihood (EL)confidence intervals/regions for the regression coefficient in a linear model with fixed design pointsand missing data. They use regression imputation method to fill in missing data, construct an ELstatistic based on‘complete sample’after imputation, and show that the EL statistic has a limitingdistribution of a weighted sum of chi-squared variables with unknown weights. They need to usean adjusted EL to obtain a confidence region on regression coefficient, in which the adjustmentcoefficient needs to be estimated. This would lead to a loss of the accuracy of the confidence re-gion. In chapter 2 of this paper, we use a new method to produce a‘complete sample’set. Basedon the data set, we construct an EL statistic which has the limiting distribution of chi-squaredvariable. Based on our result, we can construct an EL confidence region on regression coefficientwithout adjustment, which can improve the accuracy of the confidence region. Comparison ofdifference of populations is an important research topic in medical studies, economical and educa-tional fields. Qin Yongsong and Zhao Lincheng ( Semi-parametric likelihood confidence intervalsfor various differences of two populations[J], Statistics and Probability Letters, 1997, 33(2): 135-143;Empirical likelihood confidence intervals for quantile differences of two populations [J],Chinese Ann Math, 1997, 18A(6): 687-694;Semi-empirical likelihood confidence intervals forquantile differences of two samples[J], Acta Mathematicae Applicatae Sinica, 1998, 21(1): 103-112;Empirical likelihood ratio confidence intervals for various differences of two populations[J],System Science and Mathematical Sciences, 2000, 13: 23-30) systematically study the construc-tion of EL confidence intervals for various differences of two populations under complete data. Qin and Zhang ( Empirical likelihood confidence intervals for differences between two datasets withmissing data[J], Pattern Recognition Letters, 2008, 29(6):803-812) construct EL confidence in-tervals for differences of two nonparametric populations under MCAR missing mechanism. Theyuse (single) random imputation method to fill in missing data. In chapter 3 of this paper, we usefractional imputation method to impute missing data, and obtain EL confidence intervals for differ-ences of two nonparametric populations under MCAR missing mechanism, which can improve theaccuracy of the confidence intervals. In chapter 4 of this paper, we generalize the results in chapter3 to the case of MAR missing mechanism, and obtain EL confidence intervals for differences oftwo nonparametric populations under MAR missing mechanism.Here we summary some new findings in this paper.1. In studying the construction of confidence intervals for the regression coefficient in alinear model with fixed design points and missing data, we propose a new method to produce a‘complete sample’set. Based on the data set, we construct an EL statistic which has the limitingdistribution of chi-squared variable. Based on this result, we can construct an EL confidence regionon regression coefficient without adjustment, which can improve the accuracy of the confidenceregion.2. Under incomplete data and MCAR missing mechanism, we use fractional imputationmethod (a repeated imputation method) to impute missing data, and obtain EL confidence intervalsfor differences of two nonparametric populations. The usual (single) imputation method is a specialcase of fractional imputation. As the repeated time increases, fractional imputation can reduce theimputation variance. Comparing with single imputation, fractional imputation can improve theaccuracy of the confidence intervals.3. Under incomplete data and MAR missing mechanism, we use fractional imputation methodto impute missing data, and obtain EL confidence intervals for differences of two nonparametricpopulations. MAR is a weaker restriction than MCAR, and MAR is easy to be satisfied in realapplications.

  • 【分类号】O212.1
  • 【被引频次】4
  • 【下载频次】286
节点文献中: 

本文链接的文献网络图示:

本文的引文网络