节点文献

基于大语言模型的《中国小麦品种志》信息提取

Information Extraction from Chinese Wheat Varieties Journal Based on Large Language Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 韦一金陈彦清王秀东樊景超

【Author】 WEI Yijin;CHEN Yanqing;WANG Xiudong;FAN Jingchao;Agriculture Information Institution of CAAS;National Agriculture Science Data Center;Institute of Crop Sciences,CAAS;Institution of Agricultural Economics and Development,CAAS;Center for Strategic Studies,CAAS;

【通讯作者】 王秀东;

【机构】 中国农业科学院农业信息研究所国家农业科学数据中心中国农业科学院作物科学研究所中国农业科学院农业经济与发展研究所中国农业科学院战略研究中心

【摘要】 【目的】为促进小麦种质资源向小麦产业优势转化、提高小麦遗传背景丰富性,本文基于大语言模型(Large Language Model, LLM)和提示词工程,针对已出版的三卷《中国小麦品种志》进行信息挖掘。【方法】扫描《中国小麦品种志》纸质版文稿并进行OCR识别等数据处理工作以获取小麦品种数据,构建面向育种工作需求的小麦品种数据关键提取指标和相应的大语言模型提示词,以调用商业LLM api接口的方式对小麦品种数据的关键信息进行自动化提取,并形成一套成熟的基于大语言模型的小麦品种信息提取工作方案。【结果】以信息提取任务中的实际存在关系个数、识别出的关系个数、正确识别的关系个数进行精确率、召回率和F1值的计算,结果表明该小麦品种志信息提取方案在已出版的三卷《中国小麦品种志》信息提取中均达到了0.89以上的准确率、0.73以上的召回率和0.84以上的F1值。【结论】小麦品种志信息提取方案的高准确率表明其完全有能力实现精准信息提取,但是召回率又表明该方案存在部分信息无法识别的问题,因此虽然综合F1值而言该方案整体可行,但仍需对提取结果进行进一步的人工核验及审查。

【Abstract】 [Objective] In order to promote the transformation of wheat germplasm resources to wheat industry advantages and to improve the richness of wheat genetic background, this paper presents a study of information mining from the published three-volume Chinese Wheat Variety Journal based on the Large Language Model(LLM) and cue word engineering. [Methods] This project involves scanning the paper version of the Chinese Wheat Variety Journal and performing OCR recognition and other data processing tasks to obtain wheat variety data. We aim to develop key extraction indices for wheat variety data and the corresponding prompt words of the LLMs for the needs of breeding work. By calling commercial LLM API interfaces, the key information of wheat variety data will be automatically extracted. The result will be a well-established workflow for extracting wheat variety information using large language models. [Results] The calculation of precision rate, recall rate,and F1 value in terms of the number of actually existing relations, the number of recognized relations, and the number of correctly recognized relations in the information extraction task show that this wheat varietal journal information extraction scheme achieved more than 0.89 precision rate, 0.73 recall rate, and 0.84 F1 value in the information extraction for the three volumes of Chinese Wheat Varietal Journal that have been published [Conclusions] The high accuracy of this wheat varietal journal information extraction scheme indicates that it is fully capable of achieving precise information extraction, but the recall rate also indicates that the scheme has the problem that some information cannot be recognized. Though the scheme is overall feasible in terms of the combined F1 score, further manual verification and review of the extraction results is still required.

【基金】 中国农业科学院农业信息研究所科技创新工程(CAAS-ASTIP-2024-AⅡ)
  • 【文献出处】 数据与计算发展前沿(中英文) ,Frontiers of Data & Computing , 编辑部邮箱 ,2025年01期
  • 【分类号】S512.1;TP18;TP391.1
  • 【下载频次】261
节点文献中: 

本文链接的文献网络图示:

本文的引文网络