节点文献
基于SMOTE算法的女性乳腺癌风险评估模型的构建
Application of synthetic minority over-sampling technique in developing predictive model for female breast cancer
【摘要】 目的:采用合成少数类过采样技术(synthetic minority oversampling technique,SMOTE)算法处理不平衡乳腺癌筛查数据集,构建乳腺癌风险评估模型,比较基于SMOTE算法处理前后数据所建模型的准确性。方法:利用上海市2008—2012年实施的一项乳腺癌筛查项目数据,以15 046例35~74岁无乳腺癌史的上海户籍女性为研究对象,将乳腺癌筛查数据集按7:3的比例随机拆分为训练集和测试集。基于SMOTE算法进行数据重构。采用单因素logistic回归分析筛选危险因素,多因素logistic回归分析建立模型。以受试者工作特征曲线下面积(area under the curve)、敏感度、特异度、Brier评分和F1评分等为评估指标,比较基于原始数据和重构数据所建模型的预测效果。结果:年龄、初潮年龄、初产年龄、蔬菜水果日均摄入量、乳腺纤维瘤史和一级亲属乳腺癌家族史是乳腺癌的危险因素。基于SMOTE算法所建风险评估模型的AUC(95%置信区间)为0.659(0.546~0.772),优于基于原始数据建立的评估模型在测试集中的0.621(0.531~0.711)。结论:基于SMOTE算法构建的乳腺癌预测模型优于原始模型;SMOTE算法可有效解决医疗数据建模中不平衡数据的问题。
【Abstract】 Objective: To establish risk predictive models for female breast cancer using synthetic minority oversampling technique(SMOTE) algorithm to deal with unbalanced screening data, and compare the performance of established models based on original data and processed data.Methods: This analysis was based on the data derived from a breast cancer screening program conducted in 2008-2015 among 15 046 women aged 35-74 years in Shanghai. The dataset was randomly divided into a training set and a testing set by a ratio of 7:3. Logistic regression was used to identify risk factors and establish predictive models based on original data and processed data using SMOTE algorithm. Area under the curve(AUC), sensitivity, specificity, Brier score and F1 value were used to compare the performance of two established models.Results: Age, menarche age, age at first delivery, daily intake of vegetables and fruits, previous fibroadenoma of breast, and family history of breast cancer were identified as risk factors for breast cancer. The AUC and 95% confident interval of the established model based on original data was 0.621(0.531-0.711) in the testing set, while that of the model based on processed data was 0.659(0.546-0.772).Conclusion: The prediction model based on processed data using SMOTE algorithm performs better than the model based on original data. Our results suggest that SMOTE algorithm may help to solve the problem of data imbalance in medical data.
【Key words】 Breast cancer; SMOTE; Predictive model; Risk of disease occurrence;
- 【文献出处】 肿瘤 ,Tumor , 编辑部邮箱 ,2025年04期
- 【分类号】R737.9
- 【下载频次】5