Large Pre-trained Language Models (PLM) have become the most desirable starting point in the field of NLP, as they have become remarkably good at solving many individual tasks. Despite such success, in this paper, we argue that current paradigms of working with PLMs are neglecting a critical aspect of modeling human intelligence: functional compositionality. Functional compositionality - the ability to compose learned tasks - has been a long-standing challenge in the field of AI (and many other fields) as it is considered one of the hallmarks of human intelligence. An illustrative example of such is cross-lingual summarization, where a bilingual person (English-French) could directly summarize an English document into French sentences without having to translate the English document or summary into French explicitly. We discuss why this matter is an important open problem that requires further attention from the field. Then, we show that current PLMs (e.g., GPT-2 and T5) don't have functional compositionality yet and it is far from human-level generalizability. Finally, we suggest several research directions that could push the field towards zero-shot functional compositionality of language models.
翻译:大型预训练语言模型(PLM)已成为自然语言处理领域最理想的起点,因为它们在解决许多独立任务方面表现出色。尽管取得了这样的成功,但在本文中,我们认为当前使用PLM的范式忽视了模拟人类智能的一个关键方面:功能组合性。功能组合性——即组合已学习任务的能力——一直是人工智能(以及许多其他领域)领域的长期挑战,因为它被认为是人类智能的标志之一。一个说明性的例子是跨语言摘要,其中双语者(英-法)可以直接将英文文档概括为法语句子,而无需显式地将英文文档或摘要翻译成法语。我们讨论了为什么这是一个重要的开放性问题,需要该领域进一步关注。接着,我们表明当前的PLM(例如GPT-2和T5)尚不具备功能组合性,且远未达到人类水平的泛化能力。最后,我们提出了几个研究方向,有望推动该领域向语言模型的零样本功能组合性迈进。