Sentence embeddings from transformer models encode in a fixed length vector much linguistic information. We explore the hypothesis that these embeddings consist of overlapping layers of information that can be separated, and on which specific types of information -- such as information about chunks and their structural and semantic properties -- can be detected. We show that this is the case using a dataset consisting of sentences with known chunk structure, and two linguistic intelligence datasets, solving which relies on detecting chunks and their grammatical number, and respectively, their semantic roles, and through analyses of the performance on the tasks and of the internal representations built during learning.
翻译:来自Transformer模型的句子嵌入在固定长度的向量中编码了大量语言信息。我们探究了以下假设:这些嵌入由可分离的重叠信息层组成,并且可以在这些层上检测到特定类型的信息——例如关于语块及其结构和语义属性的信息。我们通过使用一个包含已知语块结构的句子数据集以及两个语言智能数据集,证明了这一假设的合理性。解决这些数据集任务依赖于检测语块及其语法数,以及分别检测其语义角色;我们通过对任务性能的分析以及学习过程中构建的内部表征的分析来验证这一点。