节点文献
中医古籍指代关系语料库的构建与指代消解方法研究
Research on Corpus Construction and Reference Resolution Methods of Ancient Literature on Traditional Chinese Medicine
【摘要】 [目的/意义]中医古籍指代消解是完整获取古籍知识、构建古籍知识图谱的关键步骤之一。然而,由于目前缺少公开发布的中医古籍指代关系语料库,导致中医古籍的指代消解研究很难开展。基于此,以中风和妊娠腹痛的古籍文献为例,构建指代关系语料库,并进行指代消解方法实验,为后续研究提供参考。[方法/过程]从《中华医典》选取中风和妊娠腹痛的古籍文献记录,定义并标注三类指代词、两类先行词和一种指代关系,人工标注构建语料库并进行质量评估。在实验阶段,指代消解被分为指称识别和指代关系预测两个子任务,实验组选用基于预训练模型BERT的指代消解模型BERT-BiLSTM-CRF,并与BiLSTM-CRF和CRF模型进行对比。[结果/结论 ]语料库标注一致性平均值达0.87,提示语料库质量良好;BERT-BiLSTM-CRF模型在指代消解任务上性能明显优于另外两种模型。总体来看,利用深度学习技术来进行中医古籍的指代消解是可行的,且利用预训练模型BERT可以更好地提高模型性能。未来的研究需要进一步增加语料库的规模,训练大模型来助力中医古籍知识发现的研究。
【Abstract】 [Purpose/Significance] The reference resolution is one of the key steps to acquire knowledge and construct the knowledge map of ancient books. However, there is a lack of published corpora of referential relations of ancient TCM literature, which makes it difficult to carry out the research of reference resolution. Based on this, this study takes ancient literature on stroke and abdominal pain during pregnancy as examples, constructs a corpus of referential relations, and conducts experiments on reference resolution methods to provide references for the following research. [Method/Process] This paper selected ancient literature records of stroke and pregnancy abdominal pain from the “Chinese Medical Code”, defined and annotated three types of pronouns, two types of antecedents and one type of referential relations, and constructed a corpus through manual annotation and evaluated its quality. In the experimental stage, the reference resolution task was divided into two subtasks: referent recognition and referential relations prediction. The experimental group selected the reference resolution model BERT-BiLSTM-CRF based on the pre-trained model BERT, and compared it with the BiLSTM-CRF and CRF models. [Result/Conclusion] The average score of corpus inter-annotator agreement(IAA) is 0.87, suggesting that the corpus quality is good. The BERT-BiLSTM-CRF model is significantly better than the other two models in the reference resolution task. In general, it is feasible to use deep learning technology to perform the reference resolution of ancient TCM literature, and the pre-trained model BERT can improve the model performance. Future research will need to further increase the scale of the corpus and train large language models to help the research of knowledge discovery on ancient TCM literature.
【Key words】 ancient TCM literature; corpus; reference relation; reference resolution;
- 【文献出处】 图书情报工作 ,Library and Information Service , 编辑部邮箱 ,2025年14期
- 【分类号】G255.1;R2-5
- 【下载频次】91