节点文献

基于LLM的等级保护事件抽取方法

LLM-based event extraction methods for classified protection

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王浩宇贾克斌谭琛瀚

【Author】 WANG Hao-yu;JIA Ke-bin;TAN Chen-han;School of Information Science and Technology,Beijing University of Technology;Beijing Laboratory of Advanced Information Network,Beijing University of Technology;

【通讯作者】 贾克斌;

【机构】 北京工业大学信息科学技术学院北京工业大学先进信息网络北京实验室

【摘要】 大语言模型在自然语言处理任务中取得了重要进展,为提高等级保护2.0语料的易读性与结构化表示,提出一个预训练后优化微调的事件抽取模型CPEE-LLM。模型基于Baichuan2-7B进行等级保护语料冻结预训练,随后构建事件抽取数据集CPEED。设计一种ChatGPT-4o提示词续写方法,生成增强数据集CPEED-ext。为优化模型效果,提出一种结果优化双重微调方法:使用ChatGPT对模型生成结果与人工抽取结果进行评估,优化数据集后再微调。实验结果显示模型在事件抽取任务中的F1值达97.7%,非等保语料剔除率达98.3%,GPT评分优于公开大模型,显著提升了等级保护语料的信息抽取效率与精准度。

【Abstract】 Significant advances have been achieved by large language models in natural language processing tasks. To enhance the readability and structured representation of Classified Protection 2. 0(CP 2. 0) corpora, an event extraction model named CPEELLM was proposed through post-pretraining optimization and fine-tuning. The model was pre-trained on classified protection corpora with frozen parameters based on Baichuan2-7B, followed by the construction of an event extraction dataset named CPEED. A ChatGPT-4o prompt-based continuation method was designed to generate an enhanced dataset, CPEED-ext. To optimize model performance, a dual fine-tuning optimization approach was introduced: model-generated outputs and manually extracted results were evaluated by ChatGPT, and the dataset was optimized before further fine-tuning. Experimental results demonstrate that the proposed model achieves an F1 score of 97. 7% in event extraction tasks, a non-CP corpus exclusion rate of 98. 3%, and outperforms public LLMs in GPT-based evaluation. These findings significantly enhance the efficiency and accuracy of information extraction for CP corpora.

【基金】 北京市自然科学基金项目(4212001)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2025年11期
  • 【分类号】TP391.1
  • 【下载频次】42
节点文献中: 

本文链接的文献网络图示:

本文的引文网络