节点文献

检索增强校验:大语言模型在重复项目识别中的应用

Gov-RAV:The Application of LLM in Duplicate Project Identification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄智勇; 张杰; 胡雪莲; 廖洁; 王瑞锦; 张凤荔; 付爽;

【Author】 HUANG Zhi-yong;ZHANG Jie;HU Xue-lian;LIAO Jie;WANG Rui-jin;ZHANG Feng-li;FU Shuang;University of Electronic Science and Technology of China;Administration for Market Regulation of Sichuan Province;

【通讯作者】 付爽;

【机构】 电子科技大学; 四川省市场监督管理局;

【摘要】 数字政府的建设一直在不断地深入推进,政务信息化项目的数量呈现指数级别的快速攀升,重复申报的现象变得越来越突出。传统的文本相似度算法受到数据孤岛以及功能重复判定领域特异性的限制约束,很难去应对重复功能判定所有的领域独特性。为了解决这个问题,文中提出了一种面向政务信息化重复项目识别的检索增强校验算法(Retrieval Augmented Verification for Government information projects,Gov-RAV)。首先,通过构建一个双粒度的项目向量库,把项目摘要和文档切片联合起来进行表征,并且设计一种智能切片算法来补全那些缺失的要素。其次,基于文本增强和大语言模型来完成功能相似度的校验工作。最后,本文利用政府的真实数据,构建了一个包含6 357条标准化记录的政务信息化功能文本对比测试集,借助这样的方式建立功能重复判定的量化基准。在这个数据集上进行的实验结果显示,Gov-RAV算法在功能重复判定方面,比向量文本相似度算法查准率高4.08个百分点,为政务信息化项目重复申报的问题提供了一种有领域适应性的智能化解决办法。

【Abstract】 As the number of government IT projects grows exponentially, duplicate submissions become increasingly prevalent. Traditional text similarity algorithms are often ineffective due to data silos and the highly domain-specific nature of functional redundancy assessment. To solve this problem, this paper proposes a retrieval-augmented verification algorithm for government information projects(Gov-RAV). The method first constructs a dual-granularity project vector database. This database integrats macro-level project summaries with micro-level document slices for joint representation. The algorithm also designs an intelligent slicing technique powered by a large language model to complete missing functional elements, significantly improving retrieval quality. Subsequently, the study introducs a function similarity verification algorithm based on text augmentation and a large language model. It builts a standardized test set containing 6 357 records from real government data to establish a quantitative benchmark. Experimental results demonstrat that the Gov-RAV algorithm achievs a 4.08 percentage point higher precision in functional duplicate identification compared to traditional vector-based text similarity methods. This work provides an adaptive and intelligent solution to the problem of duplicate submissions in government IT projects.

【基金】 国家自然科学基金(No.U2333207);四川省科技计划“揭榜挂帅”项目(No.2023YFG0374)
  • 【文献出处】 新一代信息技术 ,New Generation of Information Technology , 编辑部邮箱 ,2025年12期
  • 【分类号】TP391.3;D63
  • 【下载频次】11
节点文献中: 

本文链接的文献网络图示:

本文的引文网络