节点文献
中文重名个人名称规范记录实体匹配多元路径研究
A Multi-Path Approach to Entity Matching for Chinese Homonymous Personal Name Authority Records
【摘要】 [目的/意义]研究中文重名规范记录实体匹配的技术路径,综合运用多种算法与策略,旨在为同库或跨库同一人名称规范记录的关联和融合提供借鉴。将大语言模型与规则匹配相结合,提出重名规范记录的实体匹配方案。经数据验证发现该方案可降低中文名称规范数据库数据冗余,有助于实现数据的整合与集成。[方法/过程]选取具有较高重名率的个人名称规范记录作为分析样本,利用GPT-4o按特定规则从标目附注项和参考数据源等字段提取实体特征,通过中文大语言模型计算嵌入向量,结合最长子字符串匹配、语义距离匹配和层次聚类等技术手段,设计了4种实体匹配方案:提示词+全字段嵌入、单字段嵌入加权、最长子字符串+语义匹配计数、最长子字符串+语义距离加权。[结果/结论]实验结果表明,最长子字符串+语义匹配计数和单字段嵌入加权方法在F1值上表现突出,召回率和精确率较高,适合图书馆重名规范记录的实体匹配场景。
【Abstract】 [Purpose/Significance] By applying a range of algorithms and strategies, this study explores technical approaches of entity matching for Chinese homonymous authority records. It aims to provide a framework for the association and integration of the same individual within or across databases. By combining large language models with rule-based matching, an entity matching solution for authority records with duplicate names has been proposed. This solution is designed to reduce data redundancy and facilitate the integration of data in Chinese name authority database. [Method/Process] Personal name authority records with high repetition rate were selected as the analysis samples. Using GPT-4o, entity features were extracted from fields such as notes and reference data sources according to specific rules. Embedding vectors were calculated using a Chinese large language model. Four entity matching schemes were designed by combining techniques such as longest common substring matching, semantic distance matching, and hierarchical clustering. These schemes include: prompt + full-field embedding, single-field embedding weighting, longest common substring + semantic match count, longest common substring + semantic distance weighting. [Result/Conclusion] The experimental results show that “longest substring + semantic matching count” and the “single-field embedding weighting” method perform best in terms of F1 score, with high recall and precision. This makes them suitable for entity matching in library homonymous authority records.
【Key words】 name authority record; homonymous record; entity matching; large language model; ChatGPT;
- 【文献出处】 图书情报工作 ,Library and Information Service , 编辑部邮箱 ,2026年12期
- 【分类号】G254
- 【下载频次】78