节点文献

基于机器学习的癌症驱动框内突变预测方法研究

Research on Prediction Method of Cancer Driver Inframe Mutation Based on Machine Learning

【作者】 王震宇;

【导师】 张友华; 李军;

【作者基本信息】 安徽农业大学 , 农业硕士(专业学位), 2021, 硕士

【摘要】 随着机器学习技术不断进步,已经广泛应用于各个领域,包括计算生物学中的大数据处理、分析、预测领域,解决复杂疾病药物反应、突变预测等问题。驱动框内突变会导致癌症的发生、发展和引起癌细胞对药物反应的改变,因此我们利用机器学习方法预测癌症驱动框内突变,为揭示癌症的发生、发展机制提供帮助,为癌症的精准化药物治疗提供支持。目前,对于癌症驱动框内突变的机器学习预测工具较少,且准确率较低,输入数据的预测结果出现了较多的缺失值。为解决这一问题,本文通过对大量数据分析提取,利用计算机技术快速分析大量的生物学数据,提取数据特征。利用机器学习对癌症框内突变的预测进行研究。通过机器学习的多种方法分别建立预测模型,并利用预测结果最优的模型建立癌症驱动框内突变预测工具。提高了癌症框内突变预测的准确性。本文主要工作如下:(1)开展了基于Trans Var和Ensembl的数据筛选与标注。通过筛选收集了四个数据库的数据,作为模型的构建的训练集和测试集。利用Trans Var和ensembl注释工具查询特征工程所需要的信息,并将信息保存。然后对数据进行筛选,筛选出符合模型构建的数据。(2)提出一种多尺度下癌症数据生物学特征的提取模式。根据癌症框内突变特征提取了基因、DNA、转录本、蛋白质四个水平的7种特征。利用小提琴图分别对7个特征进行比较、分析。并对特征量化后的数据利用最大最小标准化的方法进行归一化。(3)开展了多种机器学习方法的癌症驱动框内突变预测模型研究。通过了多种机器学习算法进行模型构建,其中包括XGBoost、决策树、随机森林、Ada Boost、逻辑回归算法。并对算法模型的性能进行评估、对比。最终选择预测效果最佳的Ada Boost模型作为预测方法的模型。利用Ada Boost模型与现有VEST-Indel、DDIG、CADD三种工具进行效果对比。最终结果本文构建的Ada Boost模型在预测效果完全优与现有工具的预测效果且没有出现缺失值。(4)开展了基于Ada Boost的癌症驱动框内突变预测系统的构建。将构建完成的Ada Boost模型利用Flask框架构建成可以在线访问应用的工具,实现网站上传数据进行预测,并返回预测结果。

【Abstract】 With the continuous advancement of machine learning technology,it has been widely used in various fields,including the field of big data processing,analysis,prediction in computational biology,to solve complex disease drug response,mutation prediction and so on.Driver inframe mutations will lead to cancer initiation,progression and cause alterations in the response of cancer cells to drugs,so we used machine learning methods to predict cancer driver inframe mutations,which will provide help to reveal the mechanism of cancer initiation and progression,and provide support for precision medicine treatment of cancer.At present,there are fewer machine learning prediction tools for cancer driving inframe mutations with low accuracy and more missing values in the prediction results from the input data.To solve this problem,this paper extracts data characteristics by analyzing and extracting a large amount of biological data using computer technology quickly.Prediction of in frame mutations in cancer using machine learning.Prediction models were built separately by multiple methods of machine learning,and cancer driver inframe mutation prediction tools were built using the model with the best prediction outcome.Improves the accuracy of inframe mutation prediction in cancer.This thesis specifically works as follows:(1)Transvar and Ensembl based data filtering and annotation were undertaken..The data of the four databases were collected by filtering as the built training set and test set for the model.Information required for feature engineering is queried in annotation tools and the information saved.Data from the database were then filtered for those that fit the model building.(2)Propose a model for the extraction of biological features from cancer data at multiple scales.Seven features at the four levels of genes,DNA,transcripts,proteins were extracted based on cancer inframe mutational signatures.Violin plots were used to compare and analyze the 7 features individually.And normalization of the data after feature quantification using the method of maximum minimum normalization.(3)Several machine learning methods were developed to predict in frame mutations in cancer.Several machine learning algorithms were used for model building,including xgboost,decision tree,random forest,Ada Boost,and logistic regression.In addition,the performance of the algorithm model is evaluated and contrasted.The Ada Boost model with the best prediction performance was finally selected as the model for the prediction method.Effect contrasts were performed using Ada Boost models with three existing tools:Vest-Indel,DDIG,and CADD.Final results the Ada Boost model constructed in this paper outperformed the predictions of existing tools perfectly and did not present missing values.(4)Developed an Ada Boost based prediction system for in frame mutations in cancer drivers.The built Ada Boost model utilizes the flask framework to build into a tool that can be accessed online,enabling the website to upload data for prediction,and return prediction results.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络