节点文献

基于深度度量学习的少样本基因表达谱癌症分类

Cancer Classification Based on Deep Metric Neural Network for Low Sample Size Gene Expression Profile

【作者】 杨林

【导师】 王轩;

【作者基本信息】 哈尔滨工业大学 , 计算机科学与技术, 2020, 硕士

【摘要】 癌症已经成为全球范围内疾病和死亡的主要原因,对人们身体健康和生活都造成严重的影响。癌症产生的病因多种多样,找出癌症产生的原因和相关治疗方法成为科研工作人员的重要工作。经过科研工作者多年的研究,大多数包括癌症在内的疾病跟人类的基因相关,而人类研究自身基因的一个重要数据来源就是基因表达谱数据。基因表达数据是由生物学者选择一部分人体组织样本,加入指定试剂激活刺激组织内的基因表达,然后使用基因芯片去检测RNA蛋白质表达水平。一方面通过对某些患者及健康人士选取相同组织做基因表达谱数据,可以得到基因表达水平的不同;另一方面通过实验观测药物或者治疗方案对关键基因表达的作用,及观察前后表达水平差异,就可以评估治疗的作用和药物的疗效。因此使用基因表达谱数据对各类癌症进行细致分类,对于癌症诊断治疗有着极其重要的作用。然而在一个基因表达谱数据集中,通常只有几十个样本,而一个样本检测的基因数目高达数万,特征维度和样本数量不均衡导致基因表达谱数据直接使用机器学习模型分类时存在严重的过拟合问题。在生物信息领域使用深度学习来分析基因表达谱数据已经是一个非常重要的应用了。现有的深度学习方法已经在基于大型基因表达谱数据的癌症诊断方面取得了成功。然而,之前的深度学习模型在高维度少样本的基因表达谱数据上难以取得令人满意的表现。本文提出了一种基于深度度量学习的少样本基因表达谱数据癌症分类的方法——Deep Metric Learning with Sparse Feature Selection(DMSFS)。DMSFS通过针对基因表达谱数据高维度少样本的特点,设计基于深度度量学习的样本生成层来生成更多的新样本,从而解决样本数量和特征维度不均衡的问题。同时,DMSFS中设计了新型的基于梯度下降的特征权重层,通过模型训练中特征权重的变化幅度体现特征的重要性。将特征权重排序后,DMSFS从高维度的特征中选择重要的特征参与分类器的训练,从而减少参与训练的特征数量。DMSFS中的两个网络连接后,一方面通过样本生成层生成更多的样本促进特征权重层更好地选择重要特征,另一方面特征权重层选择更重要的特征促进模型判断样本之间差异性,从而反馈给生成层生成更适合挖掘差异性的新样本。DMSFS在与当前五个具有代表性的方法对比时,在8个真实的基因表达谱数据上实验结果取得了10至5个百分点的提高。

【Abstract】 Cancer has become a major cause of disease and death worldwide,which has a serious impact on people’s health and life.There are various causes of cancer,and it is an important task for researchers to find out the causes of cancer and relevant treatment methods.After years of research by researchers,most diseases,including cancer,are related to human genes,and an important source of data about human genes is gene expression profile data.Gene expression data are collected by biologists who select a portion of human tissue samples,add a specific reagent to activate gene expression in the tissue,and then use gene chips to detect RNA protein expression levels.On the one hand,different gene expression levels can be obtained by selecting the same tissue for gene expression profile data of some patients and healthy people.On the other hand,by observing the effects of drugs or therapeutic regimens on the expression of key genes,and the differences in expression levels before and after observation,the therapeutic effects and curative effects of drugs can be evaluated.Therefore,using gene expression profile data to analyze the key pathogenic genes of various cancers is of great importance for cancer diagnosis and treatment.However,in a data set of gene expression profile,there are usually only dozens of samples,and the number of genes detected in one sample is as high as tens of thousands.The imbalance between feature dimension and sample size leads to serious over-fitting problem when the gene expression profile data is directly classified by machine learning model.Applied deep learning to analyze gene expression profile data has been a very important application in the field of biological information.Existing deep learning methods have achieved success in cancer diagnosis based on large gene expression profile data.However,the previous deep learning model is difficult to achieve satisfy performance in high dimensional and low sample size gene expression profile data.In this thesis,we present a method for classification of cancer by gene expression based on deep metric learning——Deep Metric Learning with Sparse Feature Selection(DMSFS).DMSFS designed a new sample generation layer to generate more new samples according to the characteristics of high dimension and few samples of gene expression profile data,so as to solve the problem of unbalance between sample size and feature dimension.At the same time,a new feature weight layer is designed in DMSFS to reflect the importance of featuresthrough the change range of feature weight during model training.After ranking feature weights,DMSFS selects important features from high-dimensional features to participate in classifier for training,so as to reduce the number of features in training.After DMSFS connect the two networks,on the one hand,through the sample generation layer generates more samples for feature weight layer to choose important features,on the other hand feature weight to select the better features to get better difference through diversity loss and provide feedback to the generation layer,resulting in generate layer provides more suitable samples.DMSFS achieved an improvement of 10 to 5 percentage points on eight real gene expression profile data when compared with the current 5 representative methods.

节点文献中: