Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human's core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children's core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents' performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines.
翻译:在缺乏明确反馈的情况下,人类能够快速习得词汇的含义。儿童仅需少量被动接触即可掌握新词,这一过程被称为快速映射。这种词汇学习能力被认为是多模态理解与推理的最基础构建模块。尽管多模态学习领域近期取得进展,但针对机器实现类人词汇学习能力的系统性、严格评估仍然缺失。为填补这一空白,我们提出机器词汇学习(MEWL)基准测试,用以评估机器在具象视觉场景中学习词义的能力。MEWL涵盖人类词汇学习的核心认知工具:跨情境推理、引导式学习和语用学习。具体而言,MEWL是一个包含九项任务的少样本基准套件,用于探测不同层面的词汇学习能力。这些任务经过精心设计,与儿童词汇学习的核心能力相契合,并呼应发展心理学文献中的相关理论。通过对比分析多模态与单模态智能体与人类表现的差异,我们注意到人类与机器在词汇学习上存在显著分野。我们进一步探讨了这些差异,并呼吁机器应具备类人化的少样本词汇学习能力。