节点文献

基于深度强化学习的高频投资组合优化方法研究与实现

Research and Implementation of High-Frequency Quantitative Portfolio Optimization Method Based on Deep Reinforcement Learning

【作者】 丁南;

【导师】 韩莉;

【作者基本信息】 华东师范大学 , 电子信息(专业学位), 2024, 硕士

【摘要】 股票市场,作为企业与投资者交易的金融生态系统,截至2020年,其全球市值已超过93万亿美元。这一规模不仅为投资者创造了广泛的投资机遇,还推动了高频量化投资策略的崛起。此类策略依赖快速交易和捕捉短期市场波动来获取利润。然而,在金融市场日益复杂且竞争激烈的当下,如何有效地优化投资组合已成为投资者和金融机构亟待解决的关键问题。传统的量化投资策略主要针对低频至中频交易设计,难以应对高频环境下的复杂计算需求。现行方法大多依赖于简化的市场模型及有限的操作维度设定,在真实高频交易场景下,对复杂资产配置的适应性受限。同时当资产组合规模扩大导致行动空间维度增加时,寻求最优投资策略的难度显著加剧。为了解决高频投资组合优化中存在的问题,并应对当前市场挑战,本文致力于探讨基于深度强化学习的高频量化投资组合优化方法,并对其进行详尽的研究与实现。本文的主要工作内容如下:(1)本文创新性地提出了一种名为DRPO的基于连续决策空间的高频投资组合优化方法,旨在有效地应对上述所提及的高频投资组合优化挑战。本文通过采用马尔可夫决策过程,将投资组合优化问题转化为DRPO方法的训练目标。同时为了降低智能体交互的复杂性,在强化学习环境建模的过程中,本文设计了将状态空间分离为“静态”市场状态和“动态”资产权重状态,进一步降低了市场噪声对优化结果的影响,同时也有效地解决了高频投资组合交易中由于状态空间过于复杂导致训练时间过长的问题。(2)针对目前现有的基于深度强化学习的方法所面临的计算效率挑战,本文提出了一种概率性动态规划算法作为DRPO方法的奖励评估机制。这一方法使得智能体在训练过程中无需进行交易策略轨迹采样,进一步提升了深度强化学习智能体的效率,并获得了可靠的参数更新梯度,显著加快了算法的收敛速度,符合高频量化交易标准。(3)本文创新性地提出了一种基于多头注意力机制的投资组合优化方法RGT。RGT方法中的多头注意力机制有效地提升了Transformer模型性能,同时通过将多头注意力机制与深度确定性策略梯度DDPG结合,使得RGT方法能够通过深度强化学习的方式去建模并解决投资组合优化问题。(4)为验证本文所提出方法的可行性,本文采用了三个真实可靠的数据集进行实验,涵盖了上证50指数(A股)、道琼斯指数(美股)、加密货币三种投资资产数据,同时涵盖了多种投资组合数据集格式:LOB、OHLCV等。

【Abstract】 The stock market,functioning as a financial ecosystem where enterprises and investors transact,had exceeded a global market capitalization of 93 trillion US dollars by the end of 2020.This vast scale has not only created extensive investment opportunities for investors but also spurred the rise of high-frequency quantitative investment strategies.These strategies rely on rapid trading and the exploitation of short-term market fluctuations to generate profits.However,in today’s increasingly complex and competitive financial landscape,how to effectively optimize investment portfolios has become a critical issue that both individual investors and financial institutions are urgently seeking to address.Traditional quantitative investment strategies are primarily designed for low to mediumfrequency trading,and they struggle to cope with the complex computational demands in high-frequency environments.Contemporary approaches often rely on simplified market models and limited operational dimensions,which constrain their adaptability to intricate asset allocation scenarios within genuine high-frequency trading contexts.Concurrently,as the size of an asset portfolio expands leading to an increase in the dimensionality of the action space,the difficulty in searching for optimal investment strategies significantly intensifies.To address the issues in high-frequency portfolio optimization and respond to current market challenges,this thesis is dedicated to exploring high-frequency quantitative portfolio optimization methods based on deep reinforcement learning,conducting thorough research,and implementing them.The main contributions of this thesis are outlined as follows:(1)This thesis innovatively introduces a high-frequency portfolio optimization method named DRPO,which is based on continuous decision spaces and aims to effectively address the aforementioned challenges in high-frequency portfolio optimization.This thesis transforms the portfolio optimization problem into the training objective of the DRPO method by employing a Markov decision process.Simultaneously,to reduce the complexity of agent interaction,during the process of modeling the reinforcement learning environment,this thesis designs the separation of the state space into "static" market states and "dynamic" asset weight states.This further diminishes the impact of market noise on optimization results and effectively resolves the issue of prolonged training times in highfrequency investment portfolio trading caused by the complexity of the state space.(2)In response to the computational efficiency challenges faced by existing deep reinforcement learning-based methods,this thesis proposes a probabilistic dynamic programming algorithm as the reward evaluation mechanism of the DRPO method.This method allows the agent to train without conducting trade strategy trajectory sampling,further enhancing the efficiency of the deep reinforcement learning agent.It achieves reliable parameter update gradients,significantly accelerating the algorithm’s convergence speed,conforming to high-frequency quantitative trading standards.(3)This thesis innovatively introduces a novel portfolio optimization method called RGT,which is based on the multi-head attention mechanism.The multi-head attention mechanism in the RGT method effectively enhances the performance of the Transformer model.By combining the multi-head attention mechanism with the deep deterministic policy gradient(DDPG),the RGT method is able to model and solve investment portfolio optimization problems through deep reinforcement learning.(4)To validate the feasibility of the proposed methods in this thesis,this thesis conducts experiments using three reliable real-world datasets covering SSE 50 Index(Ashare),Dow Jones Index(U.S.stocks),and cryptocurrency assets.Additionally,it encompasses various portfolio dataset formats such as Limit Order Book(LOB)and Open,High,Low,Close,and Volume(OHLCV).

  • 【分类号】F831.51;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络