节点文献
基于多模型融合的高能宇宙线粒子鉴别
High-Energy Cosmic-Ray Particle Identification Based on Multi-model Fusion
【作者】 李洁;
【作者基本信息】 南华大学 , 应用统计(专业学位), 2025, 硕士
【摘要】 高能伽马射线的观测在宇宙物理学研究中具有重要意义。然而,由于观测数据量庞大、粒子特性复杂,传统的物理分析方法在处理效率与分类精度方面均面临较大挑战。本文围绕我国自主提出、设计并建造的新一代高能伽马射线和宇宙线地面探测装置LHAASO,探索基于机器学习的背景剔除方法,以提升伽马射线观测中的粒子分类能力。研究基于LHAASO-KM2A数据,经过数据重建、筛选、标准化及样本划分等预处理步骤,获得超过540万条事例。通过随机森林算法对各特征的重要性进行评估,选取了八个最具区分度的关键特征(Nfilt M、Base、Nu M2、Nu M3-Nu M2、Nu M4-Nu M1、Nu M1-Nu M3、Nhit M和rec_Eage),用于模型训练与评估。针对数据中存在的类别不平衡问题,采用成本敏感学习方法以提升分类模型的预测准确性和鲁棒性。本文采用了多种主流机器学习模型,包括逻辑回归、支持向量机、决策树、随机森林、XGBoost、CatBoost以及深度神经网络(DNN),并分别在四个能量区间(1012eV至1016 eV)上进行训练与测试。为进一步提高模型性能,研究将各模型的输出结果融合,通过堆叠集成算法(stacking)实现最终预测,从而在不同模型之间实现优势互补。在性能评估方面,本文采用准确率、F1分数、精确率、召回率以及接收操作特征曲线(ROC)下的面积(AUC)等多项指标进行综合评估。实验结果表明,机器学习方法显著提升了伽马射线与质子之间的分类性能。其中,XGBoost、CatBoost与DNN在分类性能上优于决策树与随机森林模型,而集成模型在背景剔除方面表现最为优异。通过集成多个模型的预测结果,有效提升了整体分类精度。模型在高能量区间的分类能力明显优于低能量区间,且评价指标值随能量的增加呈现持续上升趋势。低能量事件的分类困难主要源于伽马射线与质子在特征分布上的高度相似,以及探测信息的相对不足,导致模型识别能力受限。综上所述,本文验证了基于机器学习,特别是集成学习方法(如堆叠集成等)在高能伽马射线观测中背景剔除方面的有效性。该方法能够充分挖掘多种特征信息,在提升分类性能的同时显著增强背景排除能力,为未来大规模伽马射线观测实验提供了可靠的数据处理手段与分析技术支持,并为后续相关研究提供了可借鉴的思路与方法。
【Abstract】 The observation of high-energy gamma rays is crucial in astrophysics,yet traditional physical analysis methods struggle with large data volumes and complex particle characteristics.This study investigates a machine learning-based background rejection approach using data from LHAASO,a next-generation ground-based observatory for gamma rays and cosmic rays independently developed by China,aiming to improve particle classification performance.Based on over 5.4 million events from the LHAASO-KM2A array,data preprocessing steps—such as reconstruction,filtering,normalization,and partitioning—were performed.Feature importance was evaluated via the Random Forest algorithm,and eight key features were selected.To mitigate class imbalance,a cost-sensitive learning strategy was applied to enhance model robustness.Multiple machine learning models—including Logistic Regression,Support Vector Machine,Decision Tree,Random Forest,XGBoost,CatBoost,and Deep Neural Networks(DNN)—were trained and evaluated across four energy intervals(1012 to 1016 eV).A stacking ensemble method was employed to integrate model outputs,leveraging complementary strengths.For performance evaluation,multiple metrics were employed,including Accuracy,F1 Score,Precision,Recall,and AUC.Experimental results show that machine learning significantly improves gamma/proton classification.XGBoost,CatBoost,and DNN outperformed Decision Tree and Random Forest,with the stacking ensemble model achieving the best background suppression.Moreover,the study found that classification performance in higher energy ranges was notably better than in lower ranges,with evaluation metrics showing a consistent upward trend as energy increased.The difficulty in classifying low-energy events is primarily due to the high similarity in feature distributions between gamma rays and protons,as well as the limited information available from low-energy detections,which constrains the model’s discriminative ability.In conclusion,this study confirms the effectiveness of machine learning,particularly ensemble learning methods such as stacking,in background suppression for high-energy gamma-ray observations.It enhances classification performance through feature integration and offers a reliable framework for future experiments and related research.
【Key words】 LHAASO; cosmic ray particle; machine learning; neural networks; multi-model ensemble;
- 【网络出版投稿人】 南华大学 【网络出版年期】2026年 06期
- 【分类号】O572.1;TP18