While demographic factors like age and gender change the way people talk, and in particular, the way people talk to machines, there is little investigation into how large pre-trained language models (LMs) can adapt to these changes. To remedy this gap, we consider how demographic factors in LM language skills can be measured to determine compatibility with a target demographic. We suggest clinical techniques from Speech Language Pathology, which has norms for acquisition of language skills in humans. We conduct evaluation with a domain expert (i.e., a clinically licensed speech language pathologist), and also propose automated techniques to complement clinical evaluation at scale. Empirically, we focus on age, finding LM capability varies widely depending on task: GPT-3.5 mimics the ability of humans ranging from age 6-15 at tasks requiring inference, and simultaneously, outperforms a typical 21 year old at memorization. GPT-3.5 also has trouble with social language use, exhibiting less than 50% of the tested pragmatic skills. Findings affirm the importance of considering demographic alignment and conversational goals when using LMs as public-facing tools. Code, data, and a package will be available.
翻译:尽管年龄和性别等人统计学因素会改变人们的说话方式,尤其是人们与机器对话的方式,但目前关于大型预训练语言模型如何适应这些变化的研究尚显不足。为弥补这一空白,我们探讨了如何衡量语言模型语言技能中的人口统计学因素,以确定其与目标人群的适配性。我们借鉴了言语语言病理学的临床技术,该领域已建立人类语言技能习得规范。研究由领域专家(即持证临床言语语言病理学家)进行评估,同时提出自动化技术以补充大规模临床评估。在实证层面,我们聚焦年龄因素,发现语言模型的能力因任务不同而存在显著差异:GPT-3.5在需要推理的任务中可模拟6-15岁人类的能力,同时在记忆任务中优于典型21岁成年人的表现。GPT-3.5在社交语言使用方面存在困难,其测试语用技能达标率不足50%。研究结果证实了将语言模型作为面向公众工具使用时,考虑人口统计学适配性与对话目标的重要性。相关代码、数据集及工具包将开源提供。