Decades of research has studied how language learning infants learn to discriminate speech sounds, segment words, and associate words with their meanings. While gradual development of such capabilities is unquestionable, the exact nature of these skills and the underlying mental representations yet remains unclear. In parallel, computational studies have shown that basic comprehension of speech can be achieved by statistical learning between speech and concurrent referentially ambiguous visual input. These models can operate without prior linguistic knowledge such as representations of linguistic units, and without learning mechanisms specifically targeted at such units. This has raised the question of to what extent knowledge of linguistic units, such as phone(me)s, syllables, and words, could actually emerge as latent representations supporting the translation between speech and representations in other modalities, and without the units being proximal learning targets for the learner. In this study, we formulate this idea as the so-called latent language hypothesis (LLH), connecting linguistic representation learning to general predictive processing within and across sensory modalities. We review the extent that the audiovisual aspect of LLH is supported by the existing computational studies. We then explore LLH further in extensive learning simulations with different neural network models for audiovisual cross-situational learning, and comparing learning from both synthetic and real speech data. We investigate whether the latent representations learned by the networks reflect phonetic, syllabic, or lexical structure of input speech by utilizing an array of complementary evaluation metrics related to linguistic selectivity and temporal characteristics of the representations. As a result, we find that representations associated...
翻译:数十年的研究探讨了语言习得中的婴幼儿如何学习分辨语音、切分词汇以及将词汇与其意义关联。尽管这些能力的渐进式发展毋庸置疑,但这些技能的确切性质及其背后的心理表征仍不清晰。与此同时,计算研究表明,通过语音与同时出现的指称模糊视觉输入之间的统计学习,即可实现基本的语音理解。这些模型无需依赖语言学单位表征等先验语言知识,也无需专门针对此类单位的学习机制。这引发了一个问题:诸如音素、音节和词汇等语言学单位的知识,在何种程度上能够作为支持语音与其他模态表征之间转换的潜在表征而涌现,且无需这些单位成为学习者的直接学习目标?在本研究中,我们将这一思想定义为所谓的“潜在语言假说”(Latent Language Hypothesis, LLH),将语言学表征学习与感官模态内及跨模态的通用预测处理联系起来。我们回顾了现有计算研究对LLH视听层面的支持程度,并进一步通过不同神经网络模型的大规模学习模拟,探究基于合成语音和真实语音数据的跨情境视听学习。我们利用一系列与语言学选择性和表征时间特性相关的互补评估指标,考察网络学习的潜在表征是否反映了输入语音的音系、音节或词汇结构。结果表明,与这些表征相关的……