节点文献
基于正例和未标记样本策略及矩阵填充的miRNA相关研究
Study on miRNA with Positive and Unlabeled Learning Strategy and Matrix Completion
【作者】 王磊;
【导师】 陈兴;
【作者基本信息】 中国矿业大学 , 控制科学与工程, 2020, 硕士
【摘要】 微RNA(microRNA,miRNA)是指长约为22个核苷酸的非编码RNA,是由细胞的内源性发卡结构转录和加工而成。MiRNA常常被用来当作疾病诊断的生物标志物,而且还有研究者将miRNA当作药物的靶点进而治疗疾病。因此,发掘更多的疾病-miRNA关联将有助于理解疾病的发病机制,能促进对于疾病的诊断、预后和治疗。基于计算的预测模型可以有效地预测最可能与疾病相关的miRNA,从而降低发现新的疾病-miRNA关联的实验成本。本文所用到的数据包括疾病-miRNA关联、疾病语义相似性、miRNA功能相似性、疾病的集成相似性和miRNA的集成相似性。本文提出了两种计算模型IMCMDA和PUMDA。在IMCMDA中,疾病的集成相似性和miRNA的集成相似性作为辅助信息,并利用归纳矩阵填充算法预测潜在的疾病-miRNA关联。本文采用局部留一交叉验证、全局留一交叉验证和5折交叉验证来评估模型的性能。在IMCMDA的实验部分,本文对结肠癌、肾脏肿瘤、淋巴癌、乳腺癌和食管癌5种疾病进行案例研究。在预测的前50名疾病相关的miRNA中,分别有42、44、45、50和49个miRNA被数据库验证。在PUMDA中,先将疾病的集成相似性和miRNA的集成相似性进行拼接,得到疾病-miRNA对的特征向量。在经过特征选择后,利用有偏支持向量机(Biased Support Vector Machine,Biased-SVM)算法对所有的疾病-miRNA关联对进行预测。在PUMDA的模型性能评估部分,依旧采用局部留一交叉验证、全局留一交叉验证和5折交叉验证来评估模型的性能。在实验部分,还对食管癌、前列腺癌、肺癌和淋巴癌4种常见疾病进行案例研究,在预测的前50名疾病相关的miRNA中,分别有46、43、48和49个miRNA被数据库验证。综上所述,可知本文所提出的两种模型是有效且可靠的。
【Abstract】 MicroRNA(miRNA)is a non-coding RNA with a length about 22 nucleotides,which is transcribed and processed from the endogenous hairpin structure of cells.miRNAs are often used as biomarkers for disease diagnosis.And researchers also use miRNAs as targets of drugs for treatment.Therefore,exploring more miRNA-disease associations can help to understand the pathogenesis of disease and can promote the diagnosis,prognosis and treatment of the disease.Computation-based prediction models can effectively predict miRNAs most possible to be related to disease,thereby reducing the experimental cost of discovering new miRNA-disease associations.The data utilized in is paper include miRNA-disease association data,disease semantic similarity,miRNA functional similarity,the integrated disease similarity and the integrated miRNA similarity.Two computational methods were proposed in this paper,named IMCMDA and PUMDA,to predict potential miRNA-disease associations.In IMCMDA,we take the integrated disease similarity and the integrated miRNA similarity as side information,and then an inductive matrix completion algorithm is utilized to predict potential miRNA-disease associations.Local Leave-One-Out Cross Validation(LOOCV),global LOOCV and 5-fold cross validation were implemented to evaluate the performance of IMCMDA.In the experimental part of IMCMDA,we use IMCMDA to carry out case studies on five diseases: colon cancer,kidney cancer,lymphoma,breast cancer and esophageal cancer.Among the top 50 predicted diseaserelated miRNAs,42,44,45,50 and 49 miRNAs are verified by the databases,respectively.In PUMDA,we first stitch the integrated disease similarity and the integrated miRNA similarity to construct the feature vector of the miRNA-disease pair.After feature selection,the Biased Support Vector Machine(Biased-SVM)algorithm is adopted to predict potential miRNA-disease associations.In the part of performance evaluation,we still implement local LOOCV,global LOOCV and 5-fold cross validation to evaluate the performance of PUMDA.We also use PUMDA to implement case studies on four common diseases: esophageal cancer,prostate cancer,lung cancer and lymphoma.Among the top 50 predicted disease-related miRNAs,46,43,48 and 49 miRNAs are verified by databases,respectively.In summary,it can be seen that the two models presented in this paper are effective and reliable.
【Key words】 microRNA; disease; association prediction; matrix completion; BiasedSVM;