节点文献

基于唾液糖型的肺癌诊断模型的构建和评估及其特异性糖链的研究

Construction and Evaluation of Lung Cancer Diagnosis Model Based on Salivary Glycopatterns and Study on Its Specific N-Glycans

【作者】 张帆

【导师】 李铮;

【作者基本信息】 西北大学 , 生物化学与分子生物学, 2021, 硕士

【摘要】 研究背景:肺癌(Lung Cancer,LC)作为致死率极高的恶性疾病,其发病率在世界及中国都呈上升趋势。肺癌死亡率高的原因之一是极大部分肺癌患者最终确诊为晚期,错失最为理想的治疗时机,病理组织学检查仍是目前最为可靠的诊断依据。因此寻求非损伤性检查,进行肺癌诊断,是防控肺癌的关键。与血清样本比较,唾液采集简便且安全,无创伤,可降低血源性疾病的传播率。目前的研究工作表明在肺癌组织、患者血清和唾液中均伴有糖蛋白糖型的改变,但在患者唾液中存在的糖型变化能否用于对肺癌的诊断和对其亚型的鉴别诊断尚不清楚。本研究拟以健康志愿者(healthy volunteers,HV)、肺部良性病(benign pulmonary disease,BPD)、肺小细胞癌(smallcell lung cancer,SCLC)、肺腺癌(lung cancers with adenocarcinoma,ADC)、肺鳞癌(squamous cell carcinoma)患者唾液样本为研究对象,利用糖组学技术分析临床患者与正常对照样本中的糖蛋白糖链谱,筛选和鉴定ADC、SCC和SCLC患者唾液中共有的和特异的糖蛋白糖链,建立肺癌及其亚型的数学诊断模型并进行评估。本研究将为实现对肺癌的非损伤性诊断提供新策略。实验方法:(1)利用凝集素芯片技术对HV(50例)、BPD(47例)、ADC(54例)、SCC(43例)和SCLC(33例)患者唾液糖蛋白进行个例检测分析,以筛选出在肺癌患者中异常表达的凝集素,并利用凝集素印迹实验对结果进行验证。(2)将227例样本的凝集素芯片数据随机分为训练集与验证集,首先根据训练集数据采用二元逐步Logistic回归与ROC曲线分析分别构建Model PUD、Model LC、Model ADC和Model SCC以及Model SCLC分别用于鉴别PUD、LC、ADC、SCC以及SCLC;然后根据验证集数据通过ROC曲线分析验证上述疾病诊断模型的优劣。(3)通过决策曲线分析法(Decision Curve Analysis,DCA)对上述疾病诊断模型进行临床效益评估。(4)利用凝集素-磁性微粒复合物可特异性识别聚糖的特性分别从HV、BPD、ADC、SCC和SCLC混合唾液中分离特异糖蛋白并纯化,用PNGase F从上述糖蛋白中分离N-糖链并纯化后通过MALDI-TOF-MS进行表征。实验结果:(1)227例唾液样本凝集素芯片结果表明,相对于HV对照组,有8种凝集素(例如DBA、PHA-E和RCA120等)识别的糖链结构(例如Gal NAcα-Ser/Thr(Tn),Galβ-1,4 Glc NAc(type II),Galβ1-3Glc NAc(type I)等)在BPD、ADC和SCC以及SCLC至少一组中显著性差异表达;相对于BPD,共有4种凝集素(WFA、LTL、AAL和EEL)识别的糖链结构(例如Terminating Gal NAcα/β1-3/6Gal,Fucα1-2Galβ1-4Glc NAc,Fucα1-3(Galβ1-4)Glc NAc等)在ADC、SCC和SCLC至少一组中显著性差异表达;将HV与BPD同时作为对照组,共有12种凝集素(例如HHL、BS-I、ACA等)识别的糖链结构(例如Manα1-3Man,αGal andαGal NAc,Galβ1-3Gal NAcα-Ser/Thr(T)等)在ADC、SCC和SCLC至少一组中显著性差异表达;有6种凝集素(例如MAL-II、SNA和ECA等)识别的糖链结构(例如Siaα2-3Galβ1-4Glc(NAc)/Glc,Siaα-6Galβ1-4G lc(NAc),Galβ-1,4Glc NAc等)在BPD与HV组间具有显著性差异,而且在ADC、SCC和SCLC中也至少有一组与BPD间具有显著性差异。凝集素(随机挑选HHL、BS-I和PWM)印迹验证实验与凝集素芯片结果相符。(2)基于凝集素芯片结果,利用训练集数据构建了诊断模型Model PUD(AUC,0.903、特异性,0.875、灵敏度,0.765),Model LC(AUC,0.917、特异性,0.851、灵敏度,0.906),Model ADC(AUC,0.750、灵敏度,0.755、特异性,0.750),Model SCC(AUC,0.777、灵敏度,0.625、特异性,0.818)和Model SCLC(AUC,0.72、特异性,0.737、灵敏度,0.681)。(3)经验证集验证,Model PUD可以准确鉴别出18例HV中的13例,58例PUD患者中的56例,AUC为0.852,特异性为0.722,灵敏度为0.966,准确率为0.91;Model LC可准确鉴别出36例LC中的30例,15例BPD患者中的14例,AUC为0.881,特异性为0.833,灵敏度为0.933,准确率为0.86;Model ADC可准确鉴别出18例ADC中的14例,23例其他受试者中的17例,AUC为0.739,特异性为0.778,灵敏度为0.739,准确率为0.76;Model SCC可准确鉴别出10例SCC中的7例,31例其他受试者中的22例,AUC为0.690,特异性为0.700,灵敏度为0.710,准确率为0.71;Model SCLC,可准确鉴别出7例SCLC中的7例,25例其他受试者中的21例,AUC为0.657,特异性为0.500,灵敏度为0.840,准确率为0.72。且通过DCA分析可得,Model PUD,Model LC,Model ADC,Model SCC以及Model SCLC在绝大多数有临床效用的阈概率区间内临床效益高于其包括的单一凝集素,且Model PUD以及Model LC临床效益出色。(4)基于凝集素芯片结果可知,凝集素BS-I特异性识别的Galα1-3Gal,Galα1-6Glc,α-Gal和α-Gal Nac糖链结构在ADC、SCC和SCLC中表达水平显著升高(p<0.005)。通过MALDI-TOF/TOF-MS分析,在HV、BPD、ADC、SCC和SCLC组中分别鉴定并注释35、39、42、44和35个N-糖链,其中分别包括19、19、26、30、和23个半乳糖基化N-糖链。仅有1个半乳糖基化N-糖链(m/z 1834.635)共同存在于ADC、SCC和SCLC中。7个N-糖链(例如m/z 2498.877、2695.969、2726.988)仅存在于ADC组中,11个N-糖链(例如m/z 1995.703,2328.846,2418.853)仅存在于SCC组中,4个N-糖链(例如m/z 1647.586、1825.634、3473.273)仅存在于SCLC组中。实验结论:以上实验结论表明,相对于HV与BPD患者,肺癌患者唾液中存在异常糖基化水平改变,尤其是半乳糖基化水平显著升高且肺癌患者唾液中存在特异的半乳糖基化N-糖链。肺癌患者唾液中存在的糖型变化可用于对肺癌的诊断和对其亚型的鉴别诊断。

【Abstract】 Background: As the malignant disease with the highest mortality rate,the incidence of lung cancer(LC)is increasing year by year worldwide.One of the reasons for the high mortality of lung cancer is that the diagnosis of lung cancer is usually at an advanced stage,the optimal treatment period is missed,and pathological histology is still the gold standard for current diagnosis.Therefore,seeking non-invasive examination and performing lung cancer diagnosis is the key to the prevention and control of lung cancer.Compared with serum samples,saliva collection is safe and convenient,non-invasive,and free from the risk of blood-borne disease transmission.Current research has shown that glycoprotein glycopatterns are altered in lung cancer tissue,patient serum,and saliva,but whether glycopatterns changes present in patient saliva can be used for the diagnosis of lung cancer and the differential diagnosis of its subtypes is unknown.In this study,saliva samples from healthy volunteers(HV),patients with benign pulmonary disease(BPD),lung adenocarcinoma(ADC),lung squamous cell carcinoma(SCC),and small cell carcinoma of the lung(SCLC)were used to analyze glycan in clinical patients and normal control samples using glycomics techniques,to screen and identify common and specific glycan in saliva from ADC,SCC and SCLC patients,and to establish a mathematical diagnostic model for lung cancer and its subtypes and evaluate them.This study will provide a new strategy for achieving a non-invasive diagnosis of lung cancer.Method:(1)Salivary glycoproteins of HV(50),BPD(47),ADC(54),SCC(43)and SCLC(33)patients were detected and analyzed by lectin microarray in individual to select lectins abnormally expressed in lung cancer patients,and the results were verified by lectin blotting.(2)The lectin microarray data of 227 samples were randomly divided into retrospective and validation cohort,and Model PUD(pulmonary diseases),Model LC,Model ADC,Model SCC,and Model SCLC were firstly constructed based on the retrospective cohort data using ROC and binary stepwise logistic regression to identify PUD,LC,ADC,SCC and SCLC,respectively;The quality of the above disease diagnostic model was then verified by ROC curve analysis based on the validation cohort data.(3)the clinical benefits of the above disease diagnostic models were assessed by decision curve analysis(DCA).(4)Target glycoproteins were isolated and purified from mixed saliva of HV,BPD,ADC,SCC,and SCLC through lectin magnetic particle complexes bound to glycan structures aberrantly expressed in the saliva of lung cancer patients,respectively.Result:(1)The lectin microarray results from 227 saliva samples showed that 8 lectin(e.g.DBA,PHA-E and RCA120,etc.)recognized glycan structures(e.g.Gal NAcα-Ser/Thr(Tn),Galβ-1,4 Glc NAc(type II),Galβ1-3Glc NAc(type I),etc.)were significantly differentially expressed in at least one group of BPD,ADC and SCC and SCLC compared with HV;a total of 4 lectins(WFA,LTL,AAL,EEL)recognized glycan structures(e.g.Terminating Gal NAcα/β1-3/6Gal,Fucα1-2Galβ1-4Glc NAc,Fucα1-3(Galβ1-4)Glc NAc)were significantly differentially expressed in at least one group in ADC,SCC,SCLC relative to BPD;HV and BPD were used as the control group at the same time.A total of 12 lectins(e.g.HHL,BS-I,ACA,etc.)recognized glycan structures(e.g.Manα1-3Man,αGal andαGal NAc,Galβ1-3Gal NAcα-Ser/Thr(T))were significantly differentially expressed in at least one group of ADC,SCC,and SCLC compared with the HV and BPD groups;In addition,6 lectins(e.g.MAL-II,SNA,ECA,etc.)recognized glycan structures(e.g.Siaα2-3Galβ1-4Glc(NAc)/Glc,Siaα-6Galβ1-4G lc(NAc),Galβ-1,4Glc NAc)were significantly different between BPD and HV,and at least one of ADC,SCC,and SCLC was also significantly different from BPD.At the same time,the expression trend of randomly selected lectins HHL,BS-I,PWM blotting experimental conclusion were concordant with the lectin microarray.(2)Based on the lectin microarray results,Model PUD(AUC: 0.903、specificity: 0.875、sensitivity: 0.765),Model LC(AUC: 0.917、specificity: 0.851、sensitivity: 0.906),Model ADC(AUC: 0.750、specificity: 0.750、sensitivity: 0.755),Model SCC(AUC: 0.777、specificity: 0.818、sensitivity: 0.625),Model SCLC(AUC: 0.72、specificity: 0.737、sensitivity: 0.681)were constructed using the retrospective cohort data.(3)The validation set verified that Model PUD could accurately identify 13 of 18 HVs and56 of 58 PUD patients,with an AUC of 0.852,a specificity of 0.722,a sensitivity of 0.966,and an accuracy of 0.91;Model LC could accurately identify 30 of 36 LC and 14 of 15 BPD patients,with an AUC of 0.881,a specificity of 0.833,a sensitivity of 0.933,and an accuracy of 0.86;Model ADC could accurately identify 14 of 18 ADCs and 17 of 23 other subjects,with an AUC of 0.739,a specificity of 0.778,a sensitivity of 0.739,and an accuracy of 0.76;Model SCC could accurately identify 7 of 10 SCC and 22 of 31 other subjects,with an AUC of 0.690,a specificity of 0.700,a sensitivity of 0.710,and an accuracy of 0.71;Model SCLC,could accurately identify 7 of 7 SCLC and 21 of 25 other subjects,with an AUC of 0.657,a specificity of 0.500,a sensitivity of 0.840,and an accuracy of 0.72.And by DCA analysis,Model PUD,Model LC,Model ADC,Model SCC,and Model SCLC have higher clinical benefits than their included the single lectin included in most threshold probability intervals with clinical utility,and Model PUD and Model LC have excellent clinical benefits.(4)Based on the lectin microarray results,the glycan structures Galα1-3Gal,Galα1-6Glc,α-Gal,and α-Gal Nac specifically recognized by lectin BS-I were significantly increased in ADC,SCC,and SCLC.Available from MALDI-TOF/TOF-MS analysis,A total of 35,39,42,44,and 35 N-glycan peaks were identified and annotated in the saliva of the HV,BPD,ADC,SCC,and SCLC groups,including 19,19,26,30,and 23 galactosylated N-glycans,respectively.In total,in the three lung cancer classes of ADC,SCC,and SCLC,only galactosylated N-glycan peak(m/z 1834.635)was presented.7 peaks of N-glycan(e.g.,m/z2498.877,2695.969,2726.988)is specified in the ADC group,11 peaks of N-glycan(e.g.,m/z 1995.703,2328.846,2418.853)were unique in the SCC group,4 peaks of N-glycan(e.g.,m/z 1647.586,1825.634,3473.273)were unique in the SCLC group.Conclusion:The above findings showed that there were abnormal glycosylation level alteration in the saliva of LC relative to HV and BPD,especially the galactosylation level was significantly increased and there were specific galactosylated N-glycan in the saliva of LC.The glycopatterns changes present in the saliva of lung cancer patients can be used for the diagnosis of lung cancer and for the differential diagnosis of its subtypes.

  • 【网络出版投稿人】 西北大学
  • 【网络出版年期】2021年 12期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络