Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
翻译:预训练Transformer能够从提示中提供的示例中学习,而无需进行任何权重更新,这种非凡能力被称为上下文学习(ICL)。尽管ICL已在各个领域展现出其有效性,但其理论理解仍在发展中。现有理论大多聚焦于线性模型,而我们则在非线性回归背景下研究ICL。通过注意力中的交互机制,我们显式构建了Transformer网络以实现非线性特征,例如多项式或样条基,这些特征可张成广泛的函数类。基于这一构建,我们建立了一个框架,用于分析使用所构造特征进行的端到端上下文非线性回归。我们的理论提供了关于上下文长度和训练集大小的有限样本泛化误差界。我们通过合成回归任务对理论进行了数值验证。