Memory-based meta-learning is a technique for approximating Bayes-optimal predictors. Under fairly general conditions, minimizing sequential prediction error, measured by the log loss, leads to implicit meta-learning. The goal of this work is to investigate how far this interpretation can be realized by current sequence prediction models and training regimes. The focus is on piecewise stationary sources with unobserved switching-points, which arguably capture an important characteristic of natural language and action-observation sequences in partially observable environments. We show that various types of memory-based neural models, including Transformers, LSTMs, and RNNs can learn to accurately approximate known Bayes-optimal algorithms and behave as if performing Bayesian inference over the latent switching-points and the latent parameters governing the data distribution within each segment.
翻译:记忆元学习是一种近似贝叶斯最优预测器的技术。在相当一般的条件下,通过对数损失度量的序列预测误差最小化会导致隐式元学习。本工作的目标是探究当前序列预测模型与训练范式能在多大程度上实现这一解释。研究重点聚焦于具有未观测切换点的分段平稳源,这被认为能够捕捉部分可观测环境中自然语言与动作观测序列的重要特征。我们证明,包括Transformer、LSTM和RNN在内的多种基于记忆的神经模型能够学习准确逼近已知的贝叶斯最优算法,其行为如同对潜在切换点以及各分段内数据分布的控制参数执行贝叶斯推断。