节点文献
基于临床结构化知识的大视觉语言模型在宫颈癌放疗靶区中的勾画及亚组泛化研究
Delineation and Subgroup Generalization of Large Vision-Language Models Based on Clinical Structured Knowledge for Cervical Cancer Radiotherapy Targets
【摘要】 目的:针对宫颈癌放疗靶区自动勾画中影像与临床指南语义脱节、模型泛化能力不足等问题,构建并评估一种融合临床结构化知识的大视觉语言勾画模型在多中心、多亚组场景下的性能。方法:回顾性收集3家医疗中心478例宫颈癌患者放疗定位CT影像及临床资料,构建包含肿瘤分期、治疗方式、淋巴结转移风险及临床指南条款的结构化知识库。基于大视觉模型分割一切模型(segment anything model,SAM)架构提出K-SAM模型,通过语言编码器与跨模态注意力机制实现影像特征与指南语义深度对齐。评估K-SAM同SAM、U-Net基线模型的性能对比,并根据临床特征将患者分为盆腔早期、腹主动脉旁受累、阴道或外阴侵犯及腹主动脉旁合并腹股沟受累4个亚组进行系统评估。结果:K-SAM模型整体性能优于基线模型,其Dice相似系数(dice similarity coefficient,DSC)为0.89±0.03,95%Hausdorff距离为(5.3±0.9) mm。引入结构化知识后模型性能持续提升,从仅使用影像时DSC为0.84±0.04提高至融合全部知识后DSC为0.89±0.03,其中临床指南条款对复杂边界勾画改善最为显著。在亚组分析中,K-SAM于各亚组均保持稳定优势,其中盆腔早期组DSC为0.91±0.02,腹主动脉旁组DSC为0.87±0.03,阴道或外阴侵犯组DSC为0.89±0.03,腹主动脉旁合并腹股沟受累组DSC为0.86±0.04。结论:K-SAM通过有效融合指南语义与影像特征,提升了靶区勾画准确性与指南符合性,在多中心及复杂亚组中表现稳健,可为精准放疗提供可靠技术支持。
【Abstract】 Objective: To address the semantic gap between imaging features and clinical guidelines,as well as the insufficient generalization ability of models in automatic target delineation for cervical cancer radiotherapy,a large vision-language delineation model incorporating clinical structured knowledge was developed and evaluated for its performance in multi-center and multi-subgroup scenarios. Methods: Radiotherapy planning CT images and clinical data from 478 cervical cancer patients across 3 medical centers were retrospectively collected. A structured knowledge base incorporating tumor stage,treatment modality,lymph node metastasis risk,and clinical guideline criteria was constructed. Based on the architecture of the large vision model SAM,the K-SAM model was proposed,which achieves deep alignment between imaging features and guideline semantics via a language encoder and a cross-modal attention mechanism. The performance of K-SAM was evaluated in comparison with the SAM and U-Net baseline models. Patients were stratified into four subgroups according to clinical characteristics-pelvic early-stage disease,para-aortic involvement,vaginal or vulvar invasion,and para-aortic plus inguinal involvement-and systematically assessed. Results: The K-SAM model demonstrated superior overall performance compared to the baseline models,achieving a Dice similarity coefficient( DSC) of 0.89 ± 0.03 and a 95% Hausdorff distance of( 5.3 ±0.9) mm. Model performance improved progressively with the integration of structured knowledge,increasing from a DSC of0.84 ± 0.04( imaging only) to 0.89 ± 0.03( full knowledge integration),with clinical guideline criteria contributing most significantly to the delineation of complex boundaries. In subgroup analyses,K-SAM maintained a stable advantage across all subgroups,with a DSC of 0.91 ± 0.02 in the pelvic early-stage subgroup( PE),0.87 ± 0.03 in the para-aortic subgroup( PA),0.89 ± 0.03 in the vaginal or vulvar involvement subgroup( VV),and 0.86 ± 0.04 in the para-aortic plus inguinal involvement subgroup( PI). Conclusion: By effectively integrating guideline semantics with imaging features,the K-SAM model improves the accuracy and guideline compliance of target delineation,exhibits robust performance in multi-center settings and across complex clinical subgroups,thereby providing reliable technical support for standardized and automated precision radiotherapy.
【Key words】 Cervical cancer; Radiotherapy; Automatic target delineation; Vision-language model; Structured medical knowledge;
- 【文献出处】 肿瘤预防与治疗 ,Journal of Cancer Control and Treatment , 编辑部邮箱 ,2026年04期
- 【分类号】TP391.41;R737.33
- 【下载频次】15