In this paper, we discuss the Dutch power market, which is comprised of a day-ahead market and an intraday balancing market that operates like an auction. Due to fluctuations in power supply and demand, there is often an imbalance that leads to different prices in the two markets, providing an opportunity for arbitrage. To address this issue, we restructure the problem and propose a collaborative dual-agent reinforcement learning approach for this bi-level simulation and optimization of European power arbitrage trading. We also introduce two new implementations designed to incorporate domain-specific knowledge by imitating the trading behaviours of power traders. By utilizing reward engineering to imitate domain expertise, we are able to reform the reward system for the RL agent, which improves convergence during training and enhances overall performance. Additionally, the tranching of orders increases bidding success rates and significantly boosts profit and loss (P&L). Our study demonstrates that by leveraging domain expertise in a general learning problem, the performance can be improved substantially, and the final integrated approach leads to a three-fold improvement in cumulative P&L compared to the original agent. Furthermore, our methodology outperforms the highest benchmark policy by around 50% while maintaining efficient computational performance.
翻译:本文探讨了荷兰电力市场,该市场由日前市场和日内平衡市场(以拍卖方式运行)组成。由于电力供需波动,两个市场常出现不平衡导致价格差异,从而产生套利机会。针对这一问题,我们对问题进行重构,并提出一种协作双智能体强化学习方法,用于欧洲电力套利交易的双层仿真与优化。我们同时引入两种新实现方案,通过模仿电力交易员的交易行为来融入领域特定知识。利用奖励工程模仿领域专业知识,我们重构了强化学习智能体的奖励系统,从而改善了训练收敛性并提升了整体性能。此外,订单的分期分组提高了投标成功率,并显著提升了损益表现。研究表明,在通用学习问题中利用领域专业知识可大幅提升性能,最终集成方法相较原始智能体实现了累计损益的三倍提升。进一步地,我们的方法在保持高效计算性能的同时,比最高基准策略高出约50%。