节点文献

基于混合线性模型和条件变量分析的DNA微阵列数据分析方法研究

Study on Methods for Microarray Data Analysis Based on Mixed Linear Model Approach and Conditional Variable Analysis

【作者】 陆燕

【导师】 朱军;

【作者基本信息】 浙江大学 , 作物遗传育种, 2003, 博士

【摘要】 近年来DNA芯片技术日益成为研究大量基因表达连续变化的实验室工具。芯片技术的发展使得同时获得成千上万个基因的表达谱成为可能。DNA芯片在产生的短短几年时间已经显现出在基因序列分析、基因诊断、基因表达研究、基因组研究、发现新基因及各种病原体的诊断等生物医学领域中的应用价值。利用芯片数据,“癌变基因”的发现以及对飞速增长的基因组数据库增加功能注释等难题将迎刃而解。DNA芯片数据具有高维(成千上万个基因)和样本小(通常小于30)的特点,为了避免对伪结果进行分析,准确估计抽样方差很重要。在微阵列试验中既要包括真实的变化,又需要随机的变异。大量研究表明,聚类分析及其相关技术对于挖掘基因表达的相关模式非常有用。但是仅用这些方法不能对分析结果进行统计推断,难以得到具有生物学意义的结论,尤其是不适合分析前后时间点数据高度相关的动态基因表达数据。 本文描述的统计框架包含了基因表达分析的众多目标,与现有的分析方法完全一致,同时提高了这些方法的效用。本文着重研究差别表达基因的鉴定。本研究提出了基于混合线性模型的分析微阵列数据的方法,并将其应用于差别表达基因的鉴定、在动态或静态过程中估算基因主效应以及预测基因与环境的互作效应。用蒙特卡罗模拟对该方法的有效性和可靠性进行了比较系统的研究。这种方法可以有效地将基因表达水平根据变异来源的不同剖分为几个组成部分。主要研究内容和结论概述如下: 1.提出了分析芯片数据的一般模型,其中包括了基因、阵列效应、染料、处理效应以及基因×阵列、基因×染料、基因×处理互作效应。根据不同的试验设计,该模型可以做适当的调整。本文提出的方法主要分为两步来进行:首先,将芯片数据通过噪音过滤消除大的试验系统误差,然后在一个比较宽松的标准下通过单基因模型初步判断差异表达基因;其次,用多基因模型分析这些初定的差异表达基因以便在较严的标准下控制假阳性。用MINQUE法估计各项效应的方差和协方差分量,用AUP法预测随机效应。基因和处理的互作效应作为鉴定差异表达基因的具体指标。 2.对新提出的基于混合线性模型分析DNA芯片数据的方法用蒙特卡罗模拟进行了验证。模拟结果表明该方法在绝人多数情况下忧于传统的t检验和 WOlfinger提出的混合模型方法。验证了基因和处理的互作效应可以作为鉴定差异表达基因的更为恰当的指标。 3.研究表明我们提出的基于混合线性模型的方法可以无偏或近无偏地估算固定效应和预测随机效应。对基因主效应的无偏估计值和基因与处理互作效应的无偏预测值进行聚类可以获得具有统计学和生物学意义的结果。 4.将我们提出的混合线性模型进行拓展,可以用来分析动态的基因表达数据。我们定义了一个新变量度量给定卜1时刻的基因表达量来确定1时刻的基因表达情况,用条件变量的方法来估计条件方差、预测条件遗传效应,可以揭示在特定时间段基因表达的变异情况。 5.对新提出的基于条件变量的分析芯片数据的方法进行了蒙特卡罗模拟研究。结果表明基于条件变量的分析方法在大多数情况下表现得比差值法更有效。同时结果还进一步显示了将基囚和环境的互作效应作为鉴定差异表达基因的指标是非常有效的。 6.为了适应实际分析的需要,用C/C++语言编写了软件,可以用于分析基因芯片的表达数据,估算基因表达变异来源的方差组成和预测遗传效应,同时寻找差异表达基因。 7.以几种药物处理特异癌症细胞系的实际芯片实验数据的分析为例,说明了本研究所提方法的分析过程及分析所得结果的生物学意义。

【Abstract】 Microarrays are becoming increasingly more common laboratory tools for studying simultaneous changes in expression across a large number of genes. Recent developments in microarray technology make it possible to capture the gene expression profiles for thousands of genes at once. With this kind of data, researchers are tackling problems ranging from the identification of "cancer genes" to the formidable task of adding functional annotations to our rapidly growing gene databases. Given the high-dimensionality (thousands of genes) and small sample sizes (often <30) encountered in these datasets, an honest assessment of sampling variability is crucial and can prevent the over-interpretation of spurious results. Substantial systematic and stochastic fluctuations are involved in microarray experiments. Cluster analysis and related techniques are proving to be very useful to explore highly correlated patterns of gene expression. However, such exploratory methods alone do not provide the opportunity to engage in statistical inference and to provide results with biological sense, especially they are not fit to analyze the dynamic gene expression data which are highly correlated between time scries.We describe a statistical framework that encompasses many of the analytical goals in gene expression analysis; our framework is completely compatible with many of the current approaches and, in fact, can increase their utility. The present study has focused in the identification of differentially expressed genes in microarray data. A method for microarray data analysis based on mixed linear model approach is proposed. This method has been applied to the identification of differentially expressed genes and prediction of gene main effect and gene by environment interaction effect both in a static statement and a developmental process. Computer simulations are used to investigate the efficiency and reliability of such method under a wide range of situations. This method promises to shed light on the utilization of dividing gene expression level into several components due to various variance sources. The main results and issues are summarized as follows:1. A general genetic model for microarray data is developed, which includes effects of gene, array, dye, treatment and interaction of gene array, gene dye, and gene treatment. Proper adjustment can be made for the terms of the model according to the varied experimental project. In this paper, our method was performed in two separate steps. First, microarray data was normalized to eliminate experiment-wide systematic effects and then differentially expressed genes were prejudged under a loose standard via a single gene model. Second, these differentially expressed genes were confirmed under a stricter standard to control the false positive via a multi-gene model. Variance andcovariance components of each effect were estimated by minimum norm quadratic unbiased estimation (MINQUE) method. Adjusted unbiased p rediction (AUP) procedure w as suggested for predicting random effects. Gene treatment interaction is proposed as a measure used in identification of differentially expressed gene.2. Monte Carlo simulations are conducted to study the presented method for microarray data analysis based on mixed linear model approach. The results indicate that such method can be more effective than t-test approach and Wolfinger’s mixed model approach in a large of situations. These results have provided strong evidences suggesting gene treatment interaction as a more appropriate measure used in identifying differentially expressed genes.3. The present study has shown that our method based on mixed linear model approach can predict random effects and estimate fixed effects unbiased or asymptotic unbiased. The unbiased predicted value or estimated value of gene main effects and gene treatment interaction can be used as precursor to clustering to make sure the inputs are statistically meaningful and of biological interest.4. In the present study we extend our mixed

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2003年 02期
  • 【分类号】Q523
  • 【下载频次】329
节点文献中: