节点文献

基于机器学习算法的预后模型在肾细胞癌中的应用

Prognostic Model of Renal Cell Carcinoma Based on Machine Learning

【作者】 李浩;

【导师】 李虹;

【作者基本信息】 四川大学 , 外科(泌尿)(专业学位), 2021, 硕士

【摘要】 目的:随着肾细胞癌发病率在全球范围内持续增长,精准预测肾细胞癌患者的预后对于治疗决策的制定、个性化随访、临床试验的入组具有越来越重要的意义。目前肾细胞癌的预后模型的研究非常丰富,无论是以评分、风险分组还是Nomogram的形式,大多数研究都基于Cox比例风险模型,但这些研究均缺乏对比例风险假定及线性假定的检验。生存树及随机生存森林作为近年兴起的用于生存分析的机器学习算法,这种不依赖模型假设的非参数方法相对于传统统计学方法表现出一定的优势,同时更加适用于高维数据。本研究通过对比Cox比例风险模型、Cox扩展模型、生存树及随机生存森林模型,验证及探讨这种新型生存分析方法在临床应用中的价值。材料和方法:本研究回顾性分析了2004年至2015年监测、流行病学和结果数据库(Surveillance,Epidemiology,and End Results Program,SEER)中的肾细胞癌患者,纳入年龄、性别、种族、肿瘤分侧、肿瘤大小、TNM分期、手术方式、阳性淋巴结百分比、化疗、病理类型、Fuhrman分型等可能的预后相关因素,通过单因素及多因素Cox回归筛选出总生存率及肿瘤特异性生存率的独立预后因素,并使用Akaike’s information criterion(AIC)通过向后逐步回归法筛选出最优的模型,基于该模型建立Nomogram模型。最后使用Schoenfeld残差图法及限制性立方样条图来检验模型的比例风险假设和线性假设。通过生存数据模拟的方法建立4组生存数据,对比Cox、RCS-Cox、TD-Cox、Lasso-Cox、生存树模型、随机生存森林模型在线性、非线性、非比例风险以及高维数据4种条件下不同模型的预测精度。使用SEER肾细胞癌数据建立基于生存树算法和随机生存森林算法的预后模型,通过C-index来评价建立模型的区分度,并在The Cancer Imaging Archive(TCIA)的C4KC-Ki TS肾癌数据集中进行模型的外部验证。结果:在SEER数据库中筛选出的53505例患者中,随访时间为1~155个月,中位随访时间为53个月。诊断时中位年龄为59岁(SD=12.5),男性患者占62.6%,且男性患者(58.8岁)诊断RCC时年龄显著小于女性(60.0岁,p<0.05),病理类型以cc RCC为主(76.1%),p RCC和ch RCC分别占13.5%和5.8%。经过单因素及多因素Cox回归筛选出了总生存率及肿瘤特异生存率的独立危险因素,通过向后逐步回归法,预测总生存率的最优模型包括年龄、性别、种族、病理类型、Fuhrman分级、TNM分期、肿瘤大小、手术方式、化疗以及阳性淋巴结百分比,而预测肿瘤特异性生存率的最优模型包括年龄、病理类型、Fuhrman分级、TNM分期、肿瘤大小、手术方式、化疗以及阳性淋巴结百分比。在两个最优模型的基础上建立了预测总生存率和肿瘤特异生存率的Nomogram模型。在验证集中,Nomogram模型的区分度分别为预测总生存率的模型:0.815(95%CI:0.809–0.831),预测肿瘤特异生存率的模型:0.898(95%CI:0.882–0.911)。Schoenfeld残差检验提示模型总体违反比例风险假设(p<2-16),而限制性立方样条图发现模型中年龄及肿瘤大小两个连续型变量均存在非线性关系,当肿瘤小于5cm、年龄大于70岁时预后与肿瘤大小及年龄均表现出非线性关系。通过对比4组模拟生存数据,发现随机生存森林模型在线性(C-index=0.733,相比于Cox:C-index=0.691)、非线性(C-index=0.829,相比于RCS-Cox:C-index=0.816)、非比例风险(C-index=0.692,相对于TD-Cox:C-index=0.676)、高维数据(C-index=0.794,相对于Lasso-Cox:C-index=0.672)等情形下均相对于Cox模型及其扩展模型均表现出一定的优势。通过SEER数据建立生存树预后模型,其区分度虽然劣于Cox模型,但相较于Nomogram模型能够更加形象地对患者预后进行划分,形成代表不同预后的叶子节点。而对于随机生存森林肾癌预测模型,在验证集的区分度为OS:0.816(95%CI:0.805–0.827),CSS:0.902(95%CI:0.887–0.917),表现出不劣于Cox模型的预测效能,且能更好的拟合非线性效应。在TCIA外部验证集中,随机生存森林模型也表现出较好的泛化能力,其C-index为0.796(95%CI 0.656-0.936)。结论:常用的建立肿瘤Nomogram预后模型的方法缺少对非比例风险及非线性关系的考量,导致模型出现偏倚。本研究在生存数据模拟实验中发现随机生存森林模型在线性、非线性、非比例风险、高维数据等条件下相对于Cox模型及其扩展模型均表现出一定的优势。而生存树模型的预测能力虽然较Cox模型差,但其图形化的特征相较于Nomogram更加直观且符合决策过程。基于SEER肾癌数据建立的随机生存森林肾癌预后模型表现出不劣于Cox模型的预测效能,且能更好的拟合非线性效应,在验证集中也表现出较好地泛化能力,表明随机生存森林作为新兴的生存分析方法能够用于肾癌预后的预测。随机森林模型在高维数据的优势使其在整合多组学数据建立更加复杂模型的应用中更具前景。

【Abstract】 Objective:As the incidence of renal cell carcinoma continues to increase globally,accurate prediction of the prognosis of patients with renal cell carcinoma is of increasing significance for the formulation of treatment decisions,personalized follow-up,and the enrollment of clinical trials.At present,researches on the prognosis model of renal cell carcinoma are very rich and even redundant.Whether in the form of score,risk grouping or Nomogram,most of the studies are based on the Cox proportional hazard model,but these studies lack the test of proportional hazard assumption and linear assumption.Survival trees and random survival forests are new machine learning algorithms for survival analysis.These non-parametric methods that does not rely on model assumptions shows certain advantages over traditional statistical methods and are more suitable for high-dimensional data.This study compares the Cox proportional hazard model,the extended Cox model,survival tree,and random survival forest model,to verify and explore the value of these new survival analysis methods in clinical applications.Methods:This study retrospectively analyzed the renal cell carcinoma patients in the Surveillance,Epidemiology,and End Results Program(SEER)database from 2004 to2015.Possible prognostic factors such as age,gender,race,laterality,tumor size,TNM stage,surgical method,percentage of positive lymph nodes,chemotherapy,pathological type,Fuhrman grade were included in the study.Through univariate and multivariate Cox regression,independent prognostic factors for overall survival and tumor-specific survival were identified.Akaike’s information criterion(AIC)were used to screen out the optimal model through backward stepwise regression.Nomogram models based on the optimal model were established.Finally,Schoenfeld residual plot method and restricted cubic spline plot were used to test the proportional hazard hypothesis and linearity assumption of the model.Four sets of simulated survival data through the method of survival data simulation were generated,and then we compared the Cox model,RCS-Cox,TD-Cox,Lasso-Cox,survival tree model,random survival forest model in linear,non-linear,non-proportional hazard and high-dimensional situation of four simulated data.Then we used SEER renal cell carcinoma data to establish a prognostic model based on survival tree algorithm and random survival forest algorithm,and evaluated the discrimination ability of the established model using C-index,and compared it with SSIGN and Karakiewicz Nomogram prediction models,and in the TCIA renal cancer data set we performed external verification of the model.Results:Among the 53505 patients screened in the SEER database,the follow-up time was 1 to 155 months,and the median follow-up time was 53 months.The median age at diagnosis was 59 years old(SD = 12.5),male patients accounted for 62.6%,and male patients(58.8 years old)were significantly younger at diagnosis of RCC than females(60.0 years old,p <0.05).The pathological type was mainly cc RC(76.1%),and p RCC and ch RCC accounted for 13.5% and 5.8%,respectively.Through univariate and multivariate Cox regression,independent risk factors for overall survival and tumor-specific survival are screened out.Through backward stepwise regression,the optimal model for predicting overall survival included age,gender,race,pathological type,and Fuhrman grade,TNM staging,tumor size,surgical methods,chemotherapy,and percentage of positive lymph nodes.The best models for predicting tumor-specific survival included age,pathological type,Fuhrman grade,TNM staging,tumor size,surgical methods,chemotherapy,and positive lymph nodes percentage.On the basis of the two optimal models,Nomogram models for predicting overall survival and tumor-specific survival was established.In the validation cohort,the discrimination of Nomogram models were : predicting overall survival 0.815(95%CI: 0.809-0.831),predicting tumor-specific survival rate 0.898(95% CI: 0.882– 0.911).The Schoenfeld residual test indicated that the overall model violated the proportional hazard assumption(p <2-16),while the restricted cubic spline plot found that the two continuous variables of age and tumor size in the model have a nonlinear relationship.There is a non-linear relationship between the prognosis and tumor size and age when the tumor size less than 5cm or age older than 70 years.By comparing four sets of simulated survival data,it is found that the random survival forest model is linear(C-index = 0.733,compared to Cox,C-index = 0.691),non-linear(C-index = 0.829,compared to RCS-Cox,C-index = 0.816),non-proportional risk(C-index = 0.692,relative to TD-Cox,C-index = 0.676),high-dimensional data(C-index = 0.794,relative to Lasso-Cox,C-index = 0.672),etc.,show certain advantages over the Cox model and its extended model.The survival tree prognosis model is established through SEER data.Although its discrimination is inferior to the Cox model,it can more vividly divide the patient’s prognosis compared with the Nomogram model,forming leaf nodes that represent different prognosis.For the random survival forest kidney cancer prediction model,the discrimination in the validation set is OS: 0.816(95%CI: 0.805 – 0.827),CSS: 0.902(95%CI: 0.887 –0.917),which is not inferior to the Cox model The predictive performance of the system,and can better fit the non-linear effect.In the TCIA external validation set,the random survival forest model also showed good generalization ability,with a C-index of 0.796(95% CI 0.656-0.936).Conclusion:The commonly used method of establishing the Nomogram prognostic model of tumor lacks the consideration of non-proportional risk and nonlinear relationship,which leads to the bias of the model.In this study,the survival data simulation experiment found that the random survival forest model has certain advantages over the Cox model and its extended model under the conditions of linear,non-linear,non-proportional hazard,and high-dimensional data.Although the prediction ability of the survival tree model is worse than that of the Cox model,its graphical features are more intuitive and in line with the decision-making process than the Nomogram.The random survival forest renal cell carcinoma prognosis model established based on the SEER kidney cancer data shows no inferior predictive power to the Cox model,and can better fit nonlinear effects,and also shows good generalization ability in the validation set,indicating.As an emerging survival analysis method,random survival forest can be used to predict the prognosis of renal cell carcinoma.The advantages of random forest model in high-dimensional data make it more promising in the application of integrating multiple omics data to build more complex models.

  • 【网络出版投稿人】 四川大学
  • 【网络出版年期】2025年 02期
  • 【分类号】R737.11
节点文献中: 

本文链接的文献网络图示:

本文的引文网络