节点文献

基于样本加权的特征基因选择方法研究

Research on Feature Gene Selection Method Based on Sample Weighting

【作者】 杨丽

【导师】 陈湘涛; 李锡辉;

【作者基本信息】 湖南大学 , 计算机技术, 2013, 硕士

【摘要】 用于肿瘤诊断、预防以及治疗的分子生物标记的识别与验证是肿瘤基因组研究的重要挑战。由于临床试验以及生物验证试验需要大量的时间以及人力,因此选择一些重要的候选分子生物标记用于验证是至关重要的,而在各类肿瘤疾病研究中,基因表达谱数据被广泛用于识别候选特征基因。从机器学习的角度看,特征基因选择可以被认为是高维数据的特征选择,其目的是选择最优的少量特征子集用于解释样本表型差别,并且特征基因选择方法的鲁棒性应该较好,从而提高机器学习方法在临床诊断中的可信度。为了提高特征基因选择方法的鲁棒性并保证分类准确率,本文提出了一种基于样本加权的特征基因选择方法。该方法首先依据不同样本对于特征选择方法有不同贡献的事实,将原特征空间的样本映射到间距空间以进行样本间距分析,找出与其他样本存在显著区别的离群样本,然后为其赋予较小的权值,以减少其对特征选择方法的影响,从而提高方法的鲁棒性。然后在样本加权的基础上扩展基本过滤准则,并进行多准则融合以综合评价基因。该方法由于不仅考虑了多个准则之间的互补性,并且同时充分考虑不同样本之间的相对重要性,因此它能更全面客观地评价候选基因,从而进一步提高了算法的鲁棒性。紧接着,为了快速地搜索出较优的特征基因组合,避免基因相互作用分析中的组合爆炸问题,本文提出利用蚁群算法从筛选后的候选特征基因,结合利用多准则打分以及基因分类能力构成启发式信息,以启发式地搜索基因组合空间。最后在真实的数据集上进行比较实验,实验结果证明该方法有效保留了因为忽略样本相对重要性以及单个准则的偏袒性而被错误淘汰的有效特征基因,从而提高了特征基因子集的分类准确率,并且该方法具有更好的鲁棒性。

【Abstract】 Molecular biomarker’s identification and validation is an important challenge of the tumor genome research for tumor diagnosis, prevention, and treatment. As the clinical trials and biological verification test requires a lot of time and manpower, choosing some important biological candidate molecule numerals which are used for verification is critical, and in the study of various types of tumor diseases, the gene expression data has been widely used to identify candidate feature genes. From the machine learning perspective, gene selection can be considered as high-dimensional data of feature selection, and its purpose is to select the optimal smaller feature subset for explaining sample phenotypic differences. At the same time, the strong robustness of gene selection will enhance the enthusiastic of medical researchers. In order to improve the robustness gene selection method and to ensure the accuracy of classification, this paper presents a gene selection method based on the sample weighted.Firstly, because different sample has different contribution to the feature selection method, our method map original feature space to margin space, and then analyze the sample margin between samples. If a sample is significant different from others, its absence or existence has bigger influence to gene selection method than other samples. Thus, we can set a smaller weight to it to reduce its influence.Secondly, in order to further improve the robustness of the algorithm, we integrate multiple criteria to evaluate every gene based on sample weighted. It takes into account not only the complementarity between the multiple criteria, and at the same time can fully consider the relative importance between the samples, so that the evaluation of each gene is more objective, more comprehensive and greatly improve the robustness of the algorithm. And, in order to quickly search a better combination of genes and to avoid the combinatorial explosion of gene interaction analysis, we use ant colony algorithm to heuristic search gene combinations space.Lastly, we carry out comparison experiments on real data sets, and the experimental results show that the method is effective to retain margin effect gene which is neglect by single criterion, thereby enhancing feature gene subset classification accuracy rate, and the method has better robustness.

  • 【网络出版投稿人】 湖南大学
  • 【网络出版年期】2014年 08期
  • 【分类号】TP391.41
  • 【下载频次】90
节点文献中: