One of the most fundamental questions in quantitative finance is the existence of continuous-time diffusion models that fit market prices of a given set of options. Traditionally, one employs a mix of intuition, theoretical and empirical analysis to find models that achieve exact or approximate fits. Our contribution is to show how a suitable game theoretical formulation of this problem can help solve this question by leveraging existing developments in modern deep multi-agent reinforcement learning to search in the space of stochastic processes. Our experiments show that we are able to learn local volatility, as well as path-dependence required in the volatility process to minimize the price of a Bermudan option. Our algorithm can be seen as a particle method \textit{\`{a} la} Guyon \textit{et} Henry-Labordere where particles, instead of being designed to ensure $\sigma_{loc}(t,S_t)^2 = \mathbb{E}[\sigma_t^2|S_t]$, are learning RL-driven agents cooperating towards more general calibration targets.
翻译:量化金融中最基本的问题之一是:是否存在能够拟合给定期权集合市场价格的连续时间扩散模型。传统上,研究者通过结合直觉、理论与实证分析来寻找实现精确或近似拟合的模型。我们的贡献在于展示如何通过博弈论框架对该问题进行适当建模,并借助现代深度多智能体强化学习的最新进展,在随机过程空间中搜索解决方案。实验表明,我们的方法能够学习局部波动率以及波动率过程中所需的路径依赖性,从而最小化百慕大期权的定价误差。本算法可视为Guyon与Henry-Labordere提出的粒子法的一种推广——传统方法中粒子被设计用于确保$\sigma_{loc}(t,S_t)^2 = \mathbb{E}[\sigma_t^2|S_t]$,而我们的粒子则是由强化学习驱动的智能体,通过相互协作实现更通用的校准目标。