The objectives of option hedging/trading extend beyond mere protection against downside risks, with a desire to seek gains also driving agent's strategies. In this study, we showcase the potential of robust risk-aware reinforcement learning (RL) in mitigating the risks associated with path-dependent financial derivatives. We accomplish this by leveraging the Jaimungal, Pesenti, Wang, Tatsat (2022) and their policy gradient approach, which optimises robust risk-aware performance criteria. We specifically apply this methodology to the hedging of barrier options, and highlight how the optimal hedging strategy undergoes distortions as the agent moves from being risk-averse to risk-seeking. As well as how the agent robustifies their strategy. We further investigate the performance of the hedge when the data generating process (DGP) varies from the training DGP, and demonstrate that the robust strategies outperform the non-robust ones.
翻译:期权对冲/交易的目标不仅限于防范下行风险,追求收益同样驱动着交易主体的策略制定。本研究通过运用Jaimungal、Pesenti、Wang与Tatsat(2022)提出的策略梯度方法(该优化方法聚焦于稳健风险感知绩效准则),展示了稳健风险感知强化学习在缓解路径依赖型金融衍生品风险方面的潜力。我们专门将这一方法论应用于障碍期权的对冲场景,重点揭示了当交易主体从风险厌恶转向风险偏好时,最优对冲策略如何发生扭曲,以及主体如何强化其策略的稳健性。我们进一步考察了数据生成过程偏离训练DGP时对冲策略的表现,证实稳健策略优于非稳健策略。