We study automated intrusion response and formulate the interaction between an attacker and a defender as an optimal stopping game where attack and defense strategies evolve through reinforcement learning and self-play. The game-theoretic modeling enables us to find defender strategies that are effective against a dynamic attacker, i.e. an attacker that adapts its strategy in response to the defender strategy. Further, the optimal stopping formulation allows us to prove that optimal strategies have threshold properties. To obtain near-optimal defender strategies, we develop Threshold Fictitious Self-Play (T-FP), a fictitious self-play algorithm that learns Nash equilibria through stochastic approximation. We show that T-FP outperforms a state-of-the-art algorithm for our use case. The experimental part of this investigation includes two systems: a simulation system where defender strategies are incrementally learned and an emulation system where statistics are collected that drive simulation runs and where learned strategies are evaluated. We argue that this approach can produce effective defender strategies for a practical IT infrastructure.
翻译:我们研究自动化入侵响应问题,将攻击者与防御者之间的交互建模为最优停止博弈,其中攻击与防御策略通过强化学习和自我对弈进行演化。这种博弈论建模方法使我们能够找到对动态攻击者(即根据防御策略自适应调整自身策略的攻击者)有效的防御策略。此外,最优停止框架使我们能够证明最优策略具有阈值特性。为获取近优防御策略,我们提出了阈值虚拟自我对弈算法(T-FP),这是一种通过随机逼近学习纳什均衡的虚拟自我对弈算法。实验表明,在用例中 T-FP 的性能优于现有最优算法。本研究的实验部分包含两个系统:一个用于增量式学习防御策略的仿真系统,以及一个用于收集驱动仿真运行的统计数据并评估所学策略效果的模拟系统。我们认为该方法能够为实际IT基础设施生成有效的防御策略。