节点文献

一种索引结构优化的检索增强生成技术在保险领域的交互应用研究

A research on interactive application of index structure optimized retrieval enhanced generation technology in insurance field

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 成翌宁张正杨立马肖肖

【Author】 CHENG Yining;ZHANG Zheng;YANG Li;MA Xiaoxiao;Sinosoft Company Limited;Institute of Software Chinese Academy of Sciences;

【机构】 中科软科技股份有限公司中国科学院软件研究所

【摘要】 人工智能生成式大模型的提出对保险领域的智能交互场景产生了重大影响,在赋能行业应用软件“垂域精准计算”的技术要求的同时,为辅助代理端、业务端、用户端提供积极作用。然而大型语言模型在通用任务的生成表现中虽已经取得显著的成功,对于“垂域精准计算”面向的特定领域知识密集型任务的应用仍面临着重大限制,在处理问答即时响应时,常会产生“幻觉”现象,从而无法控制输出结果质量。仅依靠在场景应用中引入检索增强生成技术仍会存在等长切分导致上下文语义衔接被截断、相似性搜索内容过于发散检索精度缺失等痛点问题。本文提出了一种“检索增强优化索引结构的技术解决方法”,该方法在传统检索增强索引过程中增加了文档切分策略、针对块的关键词提取、语义对齐与分类、元数据补全四个技术模块,采用基于语义逻辑关系的切分方式,并基于改进的信息加权计算统计算法(term frequency-inverse document frequency, TF-IDF)实现切分段落的关键信息提取,结合引入保险行业领域词根表及业务标签库对关键词进行语义对齐、类别划分,最后完成元数据关键信息补全。在保险领域的交互应用验证结果表明,该方法有效缓解了定长切分导致语义缺失的问题,提升了知识索引结果的准确性。

【Abstract】 The proposal of artificial intelligence generative Large Language Models(LLMs) has had a significant impact on intelligent interaction scenarios in the insurance field. While empowering the technical requirements of ″vertical domain precise computing″ of industry application software, it also provides positive support for auxiliary agents, business ends, and user ends. However, despite their remarkable success in general tasks, LLMs still face significant limitations in the application of knowledge-intensive tasks in specific domains for ″vertical domain precise computing″. When dealing with real-time responses to questions and answers, ″hallucination″ phenomenon often occurs, making it impossible to control the output results. Relying solely on the introduction of retrieval enhancement generation technology in scene applications will still exist: equal length segmentation leads to the truncation of context semantic cohesion, similarity search content is too divergent and retrieval accuracy is lacking. This paper proposes a ″technical solution of retrieval enhancement optimization index structure″, which includes four technical modules in the process of retrieval enhancement index: document segmentation strategy, keyword extraction for blocks, semantic alignment and classification, metadata completion. Based on the improved information weighted calculation statistical algorithm(Term Frequency-Inverse Document Frequency, referred to as TF-IDF), the key information of segmented paragraphs is extracted, and the key words are semantically aligned and categorized by introducing the root table and business tag database in the insurance industry, and finally the key information completion of metadata is completed. The verification results of interactive application in insurance field show that this method effectively alleviates the problem of semantic loss caused by fixed-length segmentation and improves the accuracy of knowledge indexing results.

  • 【文献出处】 河北省科学院学报 ,Journal of the Hebei Academy of Sciences , 编辑部邮箱 ,2025年01期
  • 【分类号】TP391.3;F840
  • 【下载频次】43
节点文献中: 

本文链接的文献网络图示:

本文的引文网络