The objectives of option hedging/trading extend beyond mere protection against downside risks, with a desire to seek gains also driving agent's strategies. In this study, we showcase the potential of robust risk-aware reinforcement learning (RL) in mitigating the risks associated with path-dependent financial derivatives. We accomplish this by leveraging a policy gradient approach that optimises robust risk-aware performance criteria. We specifically apply this methodology to the hedging of barrier options, and highlight how the optimal hedging strategy undergoes distortions as the agent moves from being risk-averse to risk-seeking. As well as how the agent robustifies their strategy. We further investigate the performance of the hedge when the data generating process (DGP) varies from the training DGP, and demonstrate that the robust strategies outperform the non-robust ones.
翻译:期权对冲/交易的目标不仅限于防范下行风险,追求收益的动机同样驱动着交易者的策略设计。本研究展示了鲁棒风险感知强化学习在降低路径依赖型金融衍生品相关风险方面的潜力。我们通过利用一种优化鲁棒风险感知绩效标准的策略梯度方法来实现这一目标。该方法被具体应用于障碍期权的对冲,揭示了最优对冲策略如何随交易者从风险厌恶转向风险偏好而发生扭曲,以及交易者如何增强其策略的鲁棒性。我们进一步探讨了当数据生成过程偏离训练数据生成过程时对冲策略的表现,并证明鲁棒策略优于非鲁棒策略。