节点文献

PLS和SVM应用于基因表达数据分类

Partial least squares and support vector machine applied to the classification of microarray gene expression data

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 俞振超李通化吴姗

【Author】 YU Zhen-Chao, LI Tong-Hua, WU Shan(Department of Chemistry, Tongji University, Shanghai, 200092)

【机构】 同济大学化学系同济大学化学系 上海200092上海200092

【摘要】 基因表达数据的一个重要应用是给疾病样本分类,如鉴别肿瘤的类型。基因芯片的蓬勃发展使得同时测定成千上万个基因的表达成为可能。这种测定能力使得我们在很短的时间内可以得到变量数p(基因数)远远大于样本数N的数据矩阵。标准的分类统计方法在N<p的情况下通常效果不是很好,本文针对基因表达数据的特点为肿瘤的分类问题提出了一个新的分析过程。这个过程主要包括(1)通过t统计来选择基因;(2)用PCA或PLS来降维;(3)用SVM来给样本分类。在大多数情况下PLS略优于PCA,本文还给出了PLS成功预报,但PCA却预报失败的情况。最后,我们使用了重复随机学习方法来评价分类结果和方法的稳定性。

【Abstract】 One important application of microarray gene expression data is classification of diseases into categories, such as the type of tumor. With the development of gene chip, simultaneous getting thousands of genes expressions per sample come true. This ability to measure gene expression has resulted in data with the number of variables p ( genes) far exceeding the number of samples N . Standard statistical methodologies in classification and prediction do not work well or even at all when N < p. This paper propose a novel analysis procedure for classification of microarray gene expression data aiming at it’ s characteristics. The procedure involves: (1) gene selection using ( -statistics; (2) dimension reduction using Principal Component Analysis(PCA) or Partial Least Squares(PLS) ; (3) classification using Support Vector Machine( SVM). Under many circumstances PLS proves superior; this paper illustrate a condition when PCA particularly fails to predict well relative to PLS. At last, we assess the stability of classification results and methods by re-randomization studies.

【基金】 国家自然科学基金(29975019)
  • 【文献出处】 计算机与应用化学 ,Computers and Applied Chemistry , 编辑部邮箱 ,2003年05期
  • 【分类号】R-39
  • 【被引频次】21
  • 【下载频次】337
节点文献中: 

本文链接的文献网络图示:

本文的引文网络