Despite that Transformers perform well in NLP tasks, recent studies suggest that self-attention is theoretically limited in learning even some regular and context-free languages. These findings motivated us to think about their implications in modeling natural language, which is hypothesized to be mildly context-sensitive. We test Transformer's ability to learn a variety of mildly context-sensitive languages of varying complexities, and find that they generalize well to unseen in-distribution data, but their ability to extrapolate to longer strings is worse than that of LSTMs. Our analyses show that the learned self-attention patterns and representations modeled dependency relations and demonstrated counting behavior, which may have helped the models solve the languages.
翻译:尽管Transformer在自然语言处理任务中表现出色,但近期研究表明,自注意力机制在理论上甚至难以学习某些正则语言和上下文无关语言。这些发现促使我们思考其对自然语言建模的启示——自然语言被假设为轻度上下文敏感语言。我们测试了Transformer学习多种复杂度不同的轻度上下文敏感语言的能力,发现它们能很好地在训练数据分布内进行泛化,但在外推至更长字符串时的表现逊于LSTM。我们的分析表明,学习到的自注意力模式与表征能够建模依存关系并展示计数行为,这可能有助于模型解决这些语言任务。