Selective experience replay is a popular strategy for integrating lifelong learning with deep reinforcement learning. Selective experience replay aims to recount selected experiences from previous tasks to avoid catastrophic forgetting. Furthermore, selective experience replay based techniques are model agnostic and allow experiences to be shared across different models. However, storing experiences from all previous tasks make lifelong learning using selective experience replay computationally very expensive and impractical as the number of tasks increase. To that end, we propose a reward distribution-preserving coreset compression technique for compressing experience replay buffers stored for selective experience replay. We evaluated the coreset compression technique on the brain tumor segmentation (BRATS) dataset for the task of ventricle localization and on the whole-body MRI for localization of left knee cap, left kidney, right trochanter, left lung, and spleen. The coreset lifelong learning models trained on a sequence of 10 different brain MR imaging environments demonstrated excellent performance localizing the ventricle with a mean pixel error distance of 12.93 for the compression ratio of 10x. In comparison, the conventional lifelong learning model localized the ventricle with a mean pixel distance of 10.87. Similarly, the coreset lifelong learning models trained on whole-body MRI demonstrated no significant difference (p=0.28) between the 10x compressed coreset lifelong learning models and conventional lifelong learning models for all the landmarks. The mean pixel distance for the 10x compressed models across all the landmarks was 25.30, compared to 19.24 for the conventional lifelong learning models. Our results demonstrate that the potential of the coreset-based ERB compression method for compressing experiences without a significant drop in performance.
翻译:选择性经验回放是将终身学习与深度强化学习相结合的一种常用策略。其目标是通过回顾先前任务中选择的经验来避免灾难性遗忘。此外,基于选择性经验回放的技术与模型无关,且允许不同模型间共享经验。然而,随着任务数量的增加,存储所有先前任务的经验会使基于选择性经验回放的终身学习在计算上变得极其昂贵且不切实际。为此,我们提出一种保持奖励分布的核心集压缩技术,用于压缩选择性经验回放中存储的经验回放缓冲区。我们在脑肿瘤分割数据集上评估了该核心集压缩技术在脑室定位任务中的表现,并在全身MRI数据集上评估了其对左膝髌骨、左肾、右股骨转子、左肺和脾脏定位任务的效果。在连续10个不同脑部MR成像环境中训练的核心集终身学习模型,在脑室定位任务中表现出色,当压缩比为10倍时,平均像素误差距离为12.93。相比之下,传统终身学习模型的脑室定位平均像素距离为10.87。同样,在全身MRI数据集上训练的核心集终身学习模型显示,10倍压缩的核心集终身学习模型与传统终身学习模型在所有标志点定位中无显著差异(p=0.28)。10倍压缩模型在所有标志点的平均像素距离为25.30,而传统终身学习模型为19.24。我们的结果表明,基于核心集的经验回放缓冲区压缩方法在不对性能造成显著下降的情况下,具有压缩经验的潜力。