节点文献

面向乳腺癌数据的基因存储方法与拓扑数据分析

Gene Storage Method and Topological Data Analysis for Breast Cancer Data

【作者】 王瑜;

【导师】 赵毅;

【作者基本信息】 哈尔滨工业大学 , 概率论与数理统计, 2019, 硕士

【摘要】 乳腺癌是发生在乳腺腺上皮组织的恶性肿瘤,居女性恶性肿瘤的第一位,乳腺癌信息的存储和预判具有重要意义。mRNA和乳腺钼靶X线摄影成像能够对乳腺癌进行早期诊断。本文对乳腺组织数据依次做了存储、因子筛选与分类,形成一套完整的存储与分析的流程。基于乳腺癌组织和正常组织的mRNA表达水平数据和乳腺钼靶X线摄影成像,将乳腺组织信息以基因的信息存储在试管中。将数字信息转化为三进制基因编码,首尾相连地将长链基因分割为基因片段,并添加前引物、后引物和纠错位。考虑到一个信息存储试验管的“不安全性”,本文采取了分布式的存储方法。将信息存储在若干个试管中,依照同余数错位剔除信息的方法提出每个试管中的一点信息,这样只有在所有试管都存在的时候才能够恢复原始信息。添加一定的人为扰动后,通过基因池中的基因序列逐相对比可以恢复为原来的信息。进行计算机模拟,发现错误率非常低,鲁棒性高且安全性强。利用乳腺组织的开源数据集——来自不同乳腺组织的mRNA表达数据进行拓扑数据分析,用线性判别方法进行降维。构建1133维mRNA数据的单纯复形及其链复形,计算其边界算子寻找链复形的所有边缘,计算化简后的边界算子的秩的差得到了 Betti数及其拓扑特征。比较癌症组织和正常组织的Betti数和拓扑特征,寻找癌症组织的持续同调条形码大于特定参数的同调群,寻找到上述若干个同调群对应的共计53个mRNA,发现有43个能够得到文献支持,准确率高达81.13%。基于筛选出来的拓扑靶标数据,对乳腺组织做分类应用。基于决策树方法、kNN方法、随机森林方法、Adaboost方法和GBDT方法进行乳腺组织的分类,得到的准确率分别为0.975、0.75、1.0、1.0、1.0,可以发现随机森林方法、Adaboost方法、GBDT方法在mRNA表达数据集上的分类效果较佳。而深度神经网络算法由于无法有效提取一维数据的特征,分类效果不显著。本文以基因的方式存储了医疗数据,通过持续同调方法筛选得到了疾病靶标,并为早期诊断提供了依据。

【Abstract】 Breast cancer is a malignant tumor that occurs in the epithelial tissues of the breast gland.The incidence of breast cancer in the national tumor registration area ranks first among female malignant tumors.The storage and prediction of breast cancer information are of great significance.mRNA and mammography imaging can be used for early diagnosis of breast cancer.This paper has done a complete process for the storage,factor screening,and classification of breast tissue data.Based on the mRNA expression level data of breast cancer tissues and normal tissues and mammography images of mammary glands,mammary gland tissue information was stored in test tubes as genetic information.The digital information is converted into a ternary gene code,and the long-chain gene is divided into gene fragments end-to-end,and front primers,rear primers,and error correction positions are added.Considering the"insecurity" of an information storage test tube,this article adopts a distributed storage method.The information is stored in several test tubes,and a bit of information in each test tube is proposed according to the method of discarding information with congruence,so that the original information can be restored only when all the test tubes are present.After adding a certain amount of artificial disturbance,the gene sequence in the gene pool can be compared to the original information one by one.Through computer simulation,it is found that the error rate is very low,the robustness is high and the security is strong.The topological data analysis was performed using the open source data set of breast tissue—mRNA expression data from different breast tissues,and the linear discriminant method was used for dimensionality reduction.Construct a simple complex of 1133-dimensional mRNA data and its filtering simple complex,calculate its boundary operator to find all edges of the filtered simple complex,and calculate the difference between the ranks of the simplified simplified boundary operator to obtain the Betti number and its Topological characteristics.Compare the Betti number and topological characteristics of cancer tissues and normal tissues,look for homology groups with persistent homology barcodes greater than specific parameters for cancer tissues,find a total of 53 mRNAs corresponding to the above several homology groups,and find that 43 can be supported by the literature.The accuracy rate is as high as 81.13%.Based on the selected topological target data,classify and apply breast tissue.Based on decision tree method,kNN method,random forest method,Adaboost method and GBDT method for breast tissue classification,the accuracy rates obtained are 0.975,0.75,1.0,1.0,1.0 respectively.It can be found that the random forest method,Adaboost method,and GBDT method have the best classification effect on the mRNA expression dataset.However,the deep neural network algorithm can not effectively extract the characteristics of one-dimensional data,so the classification effect is not significant.In this paper,medical data is stored in a genetic manner,and disease targets are obtained through persistent homology screening.The paper provides a basis for early diagnosis.

节点文献中: