节点文献
基于区域医疗健康大数据的脑卒中发病风险预测模型的构建及验证(英文)
Development and validation of a stroke risk prediction model using regional healthcare big data and machine learning
【摘要】 目的 基于机器学习与区域医疗健康大数据构建并验证脑卒中发病风险预测模型,同时比较其与传统Logistic回归模型的预测效能。方法 采用回顾性队列研究设计,使用浙江省宁波市鄞州区2015年1月1日至2021年10月31日间的电子健康档案数据提取人口统计学特征、临床指标、生活方式、合并疾病及脑卒中家族史等信息,在单因素分析基础上,结合临床相关性和实际可操作性确定预测模型变量。将样本按7:3随机划分为训练集和验证集,分别构建传统Logistic回归模型及4种机器学习预测模型。采用混淆矩阵、受试者工作特征(receiver operating characteristic curve,ROC)曲线、校准曲线评估模型性能,并采用约登指数确定脑卒中发病风险分层的截断值。结果 共纳入92 172例研究对象,脑卒中累计发病436例,最终确定13项预测变量。混淆矩阵结果 显示,机器学习预测模型的准确率、精确率、Recall值、F1值均高于传统Logistic回归模型,其中以随机森林模型最优。ROC曲线下面积(area under the curve,AUC)显示,传统Logistic回归模型、BP神经网络、随机森林、决策树、XGBoost在训练集和验证集的AUC分别为0.777/0.779、0.921/0.918、0.988/0.980、0.980/0.955、0.962/0.958,说明机器学习预测模型性能优于传统Logistic回归模型。校准曲线显示,决策树、BP神经网络和传统Logistic回归模型与理想曲线拟合较好,优于随机森林与XGBoost模型。基于随机森林模型计算所得约登指数为0.789,表明模型截断值高于0.789被判定为高风险,反之为低风险。结论 该研究构建的机器学习预测模型性能明显优于传统Logistic回归模型且随机森林模型最优,为脑卒中高危人群的识别及分级管理提供了参考依据。
【Abstract】 Objectives: This study aimed to develop and validate a stroke risk prediction model based on machine learning(ML) and regional healthcare big data, and determine whether it may improve the prediction performance compared with the conventional Logistic Regression(LR) model.Methods: This retrospective cohort study analyzed data from the CHinese Electronic health Records Research in Yinzhou(CHERRY)(2015–2021). We included adults aged 18–75 from the platform who had established records before 2015. Individuals with pre-existing stroke, key data absence, or excessive missingness(>30 %) were excluded. Data on demographic, clinical measures, lifestyle factors, comorbidities, and family history of stroke were collected. Variable selection was performed in two stages: an initial screening via univariate analysis, followed by a prioritization of variables based on clinical relevance and actionability, with a focus on those that are modifable. Stroke prediction models were developed using LR and four ML algorithms: Decision Tree(DT), Random Forest(RF), e Xtreme Gradient Boosting(XGBoost), and Back Propagation Neural Network(BPNN). The dataset was split 7:3 for training and validation sets. Performance was assessed using receiver operating characteristic(ROC) curves,calibration, and confusion matrices, and the cutoff value was determined by Youden’s index to classify risk groups.Results: The study cohort comprised 92,172 participants with 436 incident stroke cases(incidence rate:474/100,000 person-years). Ultimately, 13 predictor variables were included. RF achieved the highest accuracy(0.935), precision(0.923), sensitivity(recall: 0.947), and F1 score(0.935). Model evaluation demonstrated superior predictive performance of ML algorithms over conventional LR, with training/validation area under the curve(AUC)s of 0.777/0.779(LR), 0.921/0.918(BPNN), 0.988/0.980(RF), 0.980/0.955(DT), and 0.962/0.958(XGBoost). Calibration analysis revealed a better ft for DT, LR and BPNN compared to RF and XGBoost model. Based on the optimal performance of the RF model, the ranking of factors in descending order of importance was: hypertension, age, diabetes, systolic blood pressure,waist, high-density lipoprotein Cholesterol, fasting blood glucose, physical activity, BMI, low-density lipoprotein cholesterol, total cholesterol, dietary habits, and family history of stroke. Using Youden’s index as the optimal cutoff, the RF model stratifed individuals into high-risk(>0.789) and low-risk(≤0.789) groups with robust discrimination.Conclusions: The ML-based prediction models demonstrated superior performance metrics compared to conventional LR and the RF is the optimal prediction model, providing an effective tool for risk stratifcation in primary stroke prevention in community settings.
【Key words】 Big data; Machine learning; Nursing; Prediction model; Stroke;
- 【文献出处】 International Journal of Nursing Sciences ,国际护理科学(英文) , 编辑部邮箱 ,2025年06期
- 【分类号】R743.3;TP311.13
- 【下载频次】60