This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.
翻译:本研究旨在探究深度强化学习作为专业执行覆盖层,能否增强高波动性加密货币市场中的配对交易。尽管该策略的传统实现在传统股票市场中被证明有效,但在应用于高方差环境时,常表现出刚性并面临严重的发散风险。为应对这一需求,本研究引入了新颖概念。为构建稳健的系统,我们开发了分层式"过滤-排序"配对选择方法,以及专有的"固定风险、自适应均值"执行模型。该系统采用含长短期记忆层的近端策略优化智能体,在严格确定性风险管理边界内管控执行决策。基于币安USD-M期货市场1小时间隔数据的评估显示,优化后的强化学习策略在样本外表现显著优于启发式基准。通过平稳循环块自助法进行的稳健性检验证实,该智能体的风险调整后超额收益在10%置信水平上具有统计显著性。尽管未达到更严格的5%阈值,这一结果凸显了数字资产极端异质方差的特征。最终,本论文通过引入统计套利与深度强化学习执行策略相结合的混合架构,为量化金融文献做出贡献。此外,它通过确定性屏蔽提供了安全强化学习的新框架,证明将神经网络策略锚定于统计稳健边界可成功缓解严重发散风险。