Recent advances in large language models have prompted researchers to examine their abilities across a variety of linguistic tasks, but little has been done to investigate how models handle the interactions in meaning across words and larger syntactic forms -- i.e. phenomena at the intersection of syntax and semantics. We present the semantic notion of agentivity as a case study for probing such interactions. We created a novel evaluation dataset by utilitizing the unique linguistic properties of a subset of optionally transitive English verbs. This dataset was used to prompt varying sizes of three model classes to see if they are sensitive to agentivity at the lexical level, and if they can appropriately employ these word-level priors given a specific syntactic context. Overall, GPT-3 text-davinci-003 performs extremely well across all experiments, outperforming all other models tested by far. In fact, the results are even better correlated with human judgements than both syntactic and semantic corpus statistics. This suggests that LMs may potentially serve as more useful tools for linguistic annotation, theory testing, and discovery than select corpora for certain tasks. Code is available at https://github.com/lindiatjuatja/lm_sem
翻译:近年来大语言模型的进展促使研究者考察其在多种语言任务中的能力,但关于模型如何处理词汇与更大句法形式之间意义交互——即句法与语义交叉现象——的研究仍相对匮乏。我们以语义概念“施事性”作为探测此类交互的案例研究。通过利用英语中一类可选及物动词的独特语言属性,我们创建了全新的评估数据集。该数据集被用于提示三种模型类的不同规模,以检验它们是否对词汇层面的施事性敏感,以及是否能在特定句法语境中适当运用这些词汇先验知识。总体而言,GPT-3 text-davinci-003在所有实验中表现极为出色,远超其他被测试模型。事实上,其结果与人类判断的相关性甚至优于句法和语义语料统计值。这表明对于特定任务,语言模型可能比精选语料更能成为语言标注、理论验证和发现的实用工具。代码参见 https://github.com/lindiatjuatja/lm_sem