节点文献

基于SMOTE算法的女性乳腺癌风险评估模型的构建

Application of synthetic minority over-sampling technique in developing predictive model for female breast cancer

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 叶尔扎提·叶尔江莫淼吴菲柳光宇许慧琳徐望红邵志敏

【Author】 YEERJIANG Yeerzhati;MO Miao;WU Fei;LIU Guangyu;XU Huilin;XU Wanghong;SHAO Zhimin;Department of Epidemiology, School of Public Health, Key Laboratory of Public Health Safety, Ministry of Education, Fudan University;Department of Breast Surgery, Fudan University Shanghai Cancer Center, Department of Oncology, Shanghai Medical College, Fudan University;Shanghai Municipal Center for Disease Control and Prevention;Center for Disease Control and Prevention of Minhang District of Shanghai;

【通讯作者】 徐望红;

【机构】 复旦大学公共卫生学院流行病学教研室,国家卫生健康委员会卫生技术评估重点实验室复旦大学附属肿瘤医院乳腺外科,复旦大学上海医学院肿瘤系上海市疾病预防控制中心上海市闵行区疾病预防控制中心

【摘要】 目的:采用合成少数类过采样技术(synthetic minority oversampling technique,SMOTE)算法处理不平衡乳腺癌筛查数据集,构建乳腺癌风险评估模型,比较基于SMOTE算法处理前后数据所建模型的准确性。方法:利用上海市2008—2012年实施的一项乳腺癌筛查项目数据,以15 046例35~74岁无乳腺癌史的上海户籍女性为研究对象,将乳腺癌筛查数据集按7:3的比例随机拆分为训练集和测试集。基于SMOTE算法进行数据重构。采用单因素logistic回归分析筛选危险因素,多因素logistic回归分析建立模型。以受试者工作特征曲线下面积(area under the curve)、敏感度、特异度、Brier评分和F1评分等为评估指标,比较基于原始数据和重构数据所建模型的预测效果。结果:年龄、初潮年龄、初产年龄、蔬菜水果日均摄入量、乳腺纤维瘤史和一级亲属乳腺癌家族史是乳腺癌的危险因素。基于SMOTE算法所建风险评估模型的AUC(95%置信区间)为0.659(0.546~0.772),优于基于原始数据建立的评估模型在测试集中的0.621(0.531~0.711)。结论:基于SMOTE算法构建的乳腺癌预测模型优于原始模型;SMOTE算法可有效解决医疗数据建模中不平衡数据的问题。

【Abstract】 Objective: To establish risk predictive models for female breast cancer using synthetic minority oversampling technique(SMOTE) algorithm to deal with unbalanced screening data, and compare the performance of established models based on original data and processed data.Methods: This analysis was based on the data derived from a breast cancer screening program conducted in 2008-2015 among 15 046 women aged 35-74 years in Shanghai. The dataset was randomly divided into a training set and a testing set by a ratio of 7:3. Logistic regression was used to identify risk factors and establish predictive models based on original data and processed data using SMOTE algorithm. Area under the curve(AUC), sensitivity, specificity, Brier score and F1 value were used to compare the performance of two established models.Results: Age, menarche age, age at first delivery, daily intake of vegetables and fruits, previous fibroadenoma of breast, and family history of breast cancer were identified as risk factors for breast cancer. The AUC and 95% confident interval of the established model based on original data was 0.621(0.531-0.711) in the testing set, while that of the model based on processed data was 0.659(0.546-0.772).Conclusion: The prediction model based on processed data using SMOTE algorithm performs better than the model based on original data. Our results suggest that SMOTE algorithm may help to solve the problem of data imbalance in medical data.

【基金】 上海市第五轮公共卫生三年行动计划(GWV-10.1-XK16);美国中华医学基金会项目(HPSS 09-991)~~
  • 【分类号】R737.9
  • 【下载频次】5
节点文献中: 

本文链接的文献网络图示:

本文的引文网络