Questions remain on the robustness of data-driven learning methods when crossing the gap from simulation to reality. We utilize weight anchoring, a method known from continual learning, to cultivate and fixate desired behavior in Neural Networks. Weight anchoring may be used to find a solution to a learning problem that is nearby the solution of another learning problem. Thereby, learning can be carried out in optimal environments without neglecting or unlearning desired behavior. We demonstrate this approach on the example of learning mixed QoS-efficient discrete resource scheduling with infrequent priority messages. Results show that this method provides performance comparable to the state of the art of augmenting a simulation environment, alongside significantly increased robustness and steerability.
翻译:关于数据驱动学习方法在跨越从仿真到现实的鸿沟时的鲁棒性问题仍然存在。我们利用持续学习中已知的权重锚定方法,来培养并固化神经网络中的期望行为。权重锚定可用于找到与另一学习问题的解邻近的一个学习问题的解。因此,学习可以在最优环境中进行,而不会忽视或遗忘期望行为。我们以混合QoS高效离散资源调度(其中包含不频繁的优先级消息)为例,展示了该方法。结果表明,该方法在提供与现有仿真环境增强方法相当的性能的同时,显著提升了鲁棒性和可操纵性。