节点文献
基于知识增强强化学习的开源信息情报推荐系统设计与实现
Design and Implementation of an Open-Source Intelligence Recommendation System Based on Knowledge-Enhanced Reinforcement Learning
【作者】 吴超;
【作者基本信息】 东南大学 , 软件工程(专业学位), 2024, 硕士
【摘要】 开源信息情报检索系统中的海量情报数据信息覆盖范围广泛、数据冗余、时效更新频繁,情报用户需要通过推荐的方式准确及时地获取满足偏好需求的情报信息。针对情报信息和检索的这些特点,本文提出基于知识增强强化学习的开源信息情报推荐方法以及其系统原型实现。其主要特点体现在通过分解推荐过程简化推荐复杂度,通过知识增强进一步表达情报的相关性和偏好性。具体工作如下:(1)分层强化学习模型的推荐复杂度简化:针对开源信息情报推荐任务候选集庞大,复杂度高,提出了一种领导层-行动层的分层强化学习推荐模型,将推荐过程分解为两个决策过程,领导层依据当前智能体状态判断对用户偏好的理解是否充分从而选择询问或推荐,询问的属性被拒绝则从情报候选集剔除关联该属性的情报以减小推荐动作空间并简化推荐任务复杂度。行动层依据领导层选项和智能体状态选择具体的属性询问或情报推荐。提出了一种内在动机机制缓解分层模型训练中领导层可能出现的稀疏奖励问题。设计了一种新的双决斗Q网络以实现分层结构下的策略学习,提出使用Gumbel-Softmax随机采样上层选项以减轻模型偏差对领导层和行动层间相互影响的不良效应。(2)知识增强的情报相关性和偏好性表达:针对情报用户偏好漂移频繁对推荐性能的影响,在强化学习智能体状态表示中融入当前时间步的知识偏好,并设计了一种基于当前偏好预测未来知识偏好的方法,以增强情报的偏好性表达,提高模型对用户偏好变化的感知。针对开源信息情报低价值密度特性对推荐性能的影响,设计了一种融合序列奖励与情报关联知识奖励的复合奖励函数。通过计算模型推荐情报序列知识嵌入与用户实际情报序列知识嵌入间的余弦差异作为知识奖励,引导模型在推荐过程中过滤低价值情报。设计了一种截断式策略梯度训练方法,旨在学习有效的知识级状态表示。与基线模型的对比实验结果显示,在开源信息情报数据集上本文模型取得了最佳推荐效果,在HR@20和NDGG@20两个指标上分别提升17.1%和2.3%,在AT指标上提升11.72%。(3)为了验证本文提出方法的可行性,本文构建了一个应用于开源信息情报推荐的原型系统。将本文方法封装成推荐算法引擎,基于Flask后端框架和Vue前端框架进行系统集成,为用户提供开源信息情报的推荐和详情展示服务。
【Abstract】 The open source intelligence retrieval system encompasses a vast array of intelligence data,characterized by wide coverage,redundancy,and frequent updates.Intelligence users require accurate and timely access to intelligence that meets their preferences through recommendation systems.In light of these characteristics of intelligence information and retrieval,this thesis proposes a knowledge-enhanced reinforcement learning-based open source intelligence recommendation method and presents its system prototype implementation.The primary features are reflected in simplifying the recommendation complexity by decomposing the recommendation process and further enhancing the expression of intelligence relevance and preferences through knowledge enhancement.The specific work is as follows:(1)Simplifying the Complexity of Hierarchical Reinforcement Learning Recommendation Model:Given the large candidate set and high complexity of open source intelligence recommendation tasks,a hierarchical reinforcement learning model comprising a leadership layer and an action layer is constructed.This model decomposes the recommendation process into two decision-making steps.The leadership layer determines whether the understanding of user preferences is sufficient based on the current agent state,thus choosing either to inquire or recommend.If an inquired attribute is rejected,intelligence associated with that attribute is excluded from the candidate set,thereby reducing the recommendation action space and simplifying the task complexity.The action layer selects specific attributes to inquire or intelligence to recommend based on the leadership layer’s option and the agent state.An intrinsic motivation mechanism is introduced to address potential sparse reward issues in the leadership layer of the hierarchical reinforcement learning model.A novel Dueling Q-Network is designed to implement policy learning in the hierarchical structure,using Gumbel-Softmax to mitigate the adverse effects of model bias on the interaction between the leadership and action layers during training.(2)Knowledge-enhanced expression of intelligence relevance and preference:To address the frequent drift in intelligence user preferences,the current timestep’s knowledge preference is integrated into the agent state representation in reinforcement learning.An inductive network is used to predict future knowledge preferences based on current preferences,thereby enhancing the expression of intelligence preferences and improving the model’s perception of user preference shifts.Given the low-value density of open source intelligence,a composite reward function integrating sequence rewards and intelligence-associated knowledge rewards is designed.The cosine difference between the knowledge embeddings of the model’s recommended intelligence sequence and the user’s actual intelligence sequence is calculated as knowledge rewards,driving the model to filter low-value intelligence during the recommendation process.The comparative experimental results with the baseline model indicate that our model achieves the best recommendation performance on the open-source information intelligence dataset.Specifically,it improves by 17.1%and 2.3%in terms of HR@20 and NDCG@20,respectively,and by 11.72%in terms of the AT metric.(3)To validate the effectiveness of the proposed method,a prototype system for OSINT recommendation is developed.The method is encapsulated into a recommendation algorithm engine,integrated with a Flask backend framework and a Vue frontend framework,providing users with recommendation services and detailed displays of open source intelligence.
【Key words】 Recommendation System; Reinforcement Learning; Knowledge Graph; Open Source Intelligence;
- 【网络出版投稿人】 东南大学 【网络出版年期】2026年 03期
- 【分类号】E91;TP391.3;TP181