节点文献
基于GWAS汇总数据的多暴露孟德尔随机化方法
Multivariable Mendelian Randomization Method Based on GWAS Summary Data
【作者】 张雪婷;
【导师】 季加东;
【作者基本信息】 山东大学 , 应用统计(专业学位), 2025, 硕士
【摘要】 孟德尔随机化(Mendelian Randomization,MR)是以遗传变异(通常是单核苷酸多态性Single Nucleotide Polymorphism,SNP)为工具变量(Instrumental Variable,Ⅳ)的因果推断方法。随全基因组关联分析(Genome-Wide Association Studies,GWAS)的发展,其广泛应用于病因推断。根据暴露数量,MR分为单暴露孟德尔随机化(Univariable Mendelian randomization,UVMR)和多暴露孟德尔随机化(Multivariable Mendelian randomization,MVMR),MVMR在暴露间存在复杂关联时可更准确估计直接因果效应。MR依赖工具变量满足相关性、排除性和独立性假设,但在实际应用中SNP往往存在水平多效性等问题,难以完全满足假设。现有的MR方法多通过SNP筛选来尽可能规避此问题,但这可能又会引发“弱工具变量偏倚”等新问题。为此,本文依据泛基因遗传效应模型,提出了基于GWAS汇总数据的多暴露孟德尔随机化方法MV-OMR(Multivariablc-Omnigcnic Mendelian Randomization)。其核心思想是,在泛基因遗传效应模型指出复杂性状是由大量SNP微小效应共同决定的前提下构建多暴露联合模型,并引入复合似然策略处理SNP间普遍存在的连锁不平衡(Linkage Disequilibrium,LD)结构,从而可纳入足够数量的SNP作为工具变量,减少筛选偏倚。此外,MV-OMR在模型中引入水平多效性项,并结合泛基因遗传效应模型设定先验结构,确保在水平多效性存在时仍能稳健估计各暴露对结局的直接因果效应。为验证MV-OMR的估计准确性和实际应用的可行性,本文进行了数值模拟和实际数据分析,并与基于全基因组工具变量思想的方法OMR以及经典的多暴露方法MVMR-IVW和MVMR-Egger进行比较分析。模拟设定了中介变量不同作用水平、不同的遗传效应以及不同的水平多效性,结果表明MV-OMR在不同情形下均能得到准确估计,特别是在中介变量存在影响时能够准确识别直接因果效应。实际数据分析中,对比MV-OMR和OMR的估计结果,发现舒张压(DBP,Diastolic Blood Pressure)和高密度脂蛋白(HDL,High-Density Lipoprotein)对冠状动脉疾病(Cardiovascular Disease,CAD)的作用效应存在差异,提示这两类因素可能通过中介变量间接影响结局,且已有研究证实了HDL的间接作用方式。未来可进一步拓展MV-OMR在更多复杂疾病和多种数据结构中的应用,增强其在病因推断中的实用性。
【Abstract】 Mendelian Randomization(MR)is a causal inference method that uses genetic variants,typically single nucleotide polymorphisms(SNPs),as instrumental variables(IVs),and has been widely used in etiological research with the rise of genome-wide association studies(GWAS).Based on exposure number,MR is classified into univariable MR(UVMR)and multivariable MR(MVMR),with the latter providing more accurate estimates in the presence of correlated exposures.However,MR relies on assumptions—relevance,exclusion restriction,and independence—that are often violated due to horizontal pleiotropy.Existing methods typically select a subset of SNPs to mitigate this,which may lead to new issues such as "weak instrument bias".To address these issues,we propose MV-OMR(Multivariable-Omnigenic Mendelian Randomization),a multivariable MR method based on the omnigenic model using GWAS summary statistics.Assuming complex traits are determined by the collective small effects of numerous SNPs,MV-OMR builds a joint exposure model and uses a composite likelihood approach to account for widespread linkage disequilibrium(LD),allowing more SNPs as IVs and reducing selection bias.It also models horizontal pleiotropy explicitly and applies prior structures from the omnigenic model,ensuring robust estimation of direct effects.To assess the estimation accuracy and practical applicability of MV-OMR,we conducted simulation studies and real data analyses,benchmarking its performance against both the genome-wide IV-based OMR approach and established multivariable methods(MVMR-IVW and MVMR-Egger).Simulations were designed under varying mediation effects,genetic architectures,and levels of horizontal pleiotropy.Results demonstrated that MV-OMR consistently produced accurate estimates,particularly when mediators were present,effectively identifying direct causal effects.When comparing between MV-OMR and OMR eatimates in the real data analysis,we revealed differences in the estimated effects of diastolic blood pressure(DBP)and high-density lipoprotein(HDL)on coronary artery disease(CAD),suggesting these factors may influence CAD indirectly through mediators.This is supported by previous findings on HDL’s indirect effect.MV-OMR shows strong potential for broader application in complex disease research and diverse data settings,improving its value in causal inference.
【Key words】 Multivariable Mendelian Randomization; the Omnigenic Model; Linkage Disequilibrium; Composite likelihood;
- 【网络出版投稿人】 山东大学 【网络出版年期】2026年 06期
- 【分类号】O212.1;Q811.4