Human infant learning happens during exploration of the environment, by interaction with objects, and by listening to and repeating utterances casually, which is analogous to unsupervised learning. Only occasionally, a learning infant would receive a matching verbal description of an action it is committing, which is similar to supervised learning. Such a learning mechanism can be mimicked with deep learning. We model this weakly supervised learning paradigm using our Paired Gated Autoencoders (PGAE) model, which combines an action and a language autoencoder. After observing a performance drop when reducing the proportion of supervised training, we introduce the Paired Transformed Autoencoders (PTAE) model, using Transformer-based crossmodal attention. PTAE achieves significantly higher accuracy in language-to-action and action-to-language translations, particularly in realistic but difficult cases when only few supervised training samples are available. We also test whether the trained model behaves realistically with conflicting multimodal input. In accordance with the concept of incongruence in psychology, conflict deteriorates the model output. Conflicting action input has a more severe impact than conflicting language input, and more conflicting features lead to larger interference. PTAE can be trained on mostly unlabelled data where labeled data is scarce, and it behaves plausibly when tested with incongruent input.
翻译:人类婴儿的学习发生在探索环境、与物体互动、随意聆听和重复话语的过程中,这类似于无监督学习。仅偶尔情况下,学习中的婴儿会收到对其正在执行动作的匹配语言描述,这类似于监督学习。这种学习机制可通过深度学习进行模拟。我们利用配对门控自编码器(Paired Gated Autoencoders, PGAE)模型对这种弱监督学习范式进行建模,该模型结合了动作自编码器和语言自编码器。在观察到减少监督训练比例会导致性能下降后,我们引入了基于Transformer跨模态注意力的配对变换自编码器(Paired Transformed Autoencoders, PTAE)模型。PTAE在语言到动作和动作到语言的翻译中实现了显著更高的准确率,尤其是在仅有少量监督训练样本的现实困难场景下。我们还测试了训练模型在面对冲突的多模态输入时的表现是否符合现实。根据心理学中的不协调概念,冲突会劣化模型输出:冲突性动作输入的影响比冲突性语言输入更为严重,且冲突特征越多,干扰越大。PTAE可在标注数据稀缺时主要利用未标注数据进行训练,并在接受冲突输入测试时表现出合理的响应行为。