节点文献
基于大语言模型的《中国小麦品种志》信息提取
Information Extraction from Chinese Wheat Varieties Journal Based on Large Language Model
【摘要】 【目的】为促进小麦种质资源向小麦产业优势转化、提高小麦遗传背景丰富性,本文基于大语言模型(Large Language Model, LLM)和提示词工程,针对已出版的三卷《中国小麦品种志》进行信息挖掘。【方法】扫描《中国小麦品种志》纸质版文稿并进行OCR识别等数据处理工作以获取小麦品种数据,构建面向育种工作需求的小麦品种数据关键提取指标和相应的大语言模型提示词,以调用商业LLM api接口的方式对小麦品种数据的关键信息进行自动化提取,并形成一套成熟的基于大语言模型的小麦品种信息提取工作方案。【结果】以信息提取任务中的实际存在关系个数、识别出的关系个数、正确识别的关系个数进行精确率、召回率和F1值的计算,结果表明该小麦品种志信息提取方案在已出版的三卷《中国小麦品种志》信息提取中均达到了0.89以上的准确率、0.73以上的召回率和0.84以上的F1值。【结论】小麦品种志信息提取方案的高准确率表明其完全有能力实现精准信息提取,但是召回率又表明该方案存在部分信息无法识别的问题,因此虽然综合F1值而言该方案整体可行,但仍需对提取结果进行进一步的人工核验及审查。
【Abstract】 [Objective] In order to promote the transformation of wheat germplasm resources to wheat industry advantages and to improve the richness of wheat genetic background, this paper presents a study of information mining from the published three-volume Chinese Wheat Variety Journal based on the Large Language Model(LLM) and cue word engineering. [Methods] This project involves scanning the paper version of the Chinese Wheat Variety Journal and performing OCR recognition and other data processing tasks to obtain wheat variety data. We aim to develop key extraction indices for wheat variety data and the corresponding prompt words of the LLMs for the needs of breeding work. By calling commercial LLM API interfaces, the key information of wheat variety data will be automatically extracted. The result will be a well-established workflow for extracting wheat variety information using large language models. [Results] The calculation of precision rate, recall rate,and F1 value in terms of the number of actually existing relations, the number of recognized relations, and the number of correctly recognized relations in the information extraction task show that this wheat varietal journal information extraction scheme achieved more than 0.89 precision rate, 0.73 recall rate, and 0.84 F1 value in the information extraction for the three volumes of Chinese Wheat Varietal Journal that have been published [Conclusions] The high accuracy of this wheat varietal journal information extraction scheme indicates that it is fully capable of achieving precise information extraction, but the recall rate also indicates that the scheme has the problem that some information cannot be recognized. Though the scheme is overall feasible in terms of the combined F1 score, further manual verification and review of the extraction results is still required.
【Key words】 Large Language Model(LLM); agriculture; wheat; information mining; genetic resources;
- 【文献出处】 数据与计算发展前沿(中英文) ,Frontiers of Data & Computing , 编辑部邮箱 ,2025年01期
- 【分类号】S512.1;TP18;TP391.1
- 【下载频次】261