The traveling officer problem (TOP) is a challenging stochastic optimization task. In this problem, a parking officer is guided through a city equipped with parking sensors to fine as many parking offenders as possible. A major challenge in TOP is the dynamic nature of parking offenses, which randomly appear and disappear after some time, regardless of whether they have been fined. Thus, solutions need to dynamically adjust to currently fineable parking offenses while also planning ahead to increase the likelihood that the officer arrives during the offense taking place. Though various solutions exist, these methods often struggle to take the implications of actions on the ability to fine future parking violations into account. This paper proposes SATOP, a novel spatial-aware deep reinforcement learning approach for TOP. Our novel state encoder creates a representation of each action, leveraging the spatial relationships between parking spots, the agent, and the action. Furthermore, we propose a novel message-passing module for learning future inter-action correlations in the given environment. Thus, the agent can estimate the potential to fine further parking violations after executing an action. We evaluate our method using an environment based on real-world data from Melbourne. Our results show that SATOP consistently outperforms state-of-the-art TOP agents and is able to fine up to 22% more parking offenses.
翻译:巡逻官员问题(TOP)是一项具有挑战性的随机优化任务。在该问题中,停车巡逻官员需借助城市中部署的停车传感器进行导航,以尽可能多地处罚违章停车行为。TOP的主要挑战在于停车违章行为的动态性——这些行为会随机出现,并在一定时间后自行消失(无论是否已被处罚)。因此,解决方案需要动态调整以适应当前可处罚的违章行为,同时前瞻性地规划路径,提高巡逻官员在违章持续期间到达现场的概率。尽管已有多种解决方案,但这些方法通常难以充分评估当前行动对未来处罚违章停车能力的影响。本文提出SATOP,一种面向TOP的新型空间感知深度强化学习方法。我们的新型状态编码器为每个动作创建表征,充分利用停车位、智能体与动作之间的空间关系。此外,我们提出了一种新颖的消息传递模块,用于学习给定环境中动作间的未来相关性。因此,智能体能够估计执行某个动作后进一步处罚违章停车的潜力。我们基于墨尔本真实数据构建环境进行方法评估。结果表明,SATOP持续优于当前最先进的TOP智能体,可多处罚高达22%的违章停车行为。