The acquisition of grammar has been a central question to adjudicate between theories of language acquisition. In order to conduct faster, more reproducible, and larger-scale corpus studies on grammaticality in child-caregiver conversations, tools for automatic annotation can offer an effective alternative to tedious manual annotation. We propose a coding scheme for context-dependent grammaticality in child-caregiver conversations and annotate more than 4,000 utterances from a large corpus of transcribed conversations. Based on these annotations, we train and evaluate a range of NLP models. Our results show that fine-tuned Transformer-based models perform best, achieving human inter-annotation agreement levels.As a first application and sanity check of this tool, we use the trained models to annotate a corpus almost two orders of magnitude larger than the manually annotated data and verify that children's grammaticality shows a steady increase with age.This work contributes to the growing literature on applying state-of-the-art NLP methods to help study child language acquisition at scale.
翻译:语法习得一直是判断语言习得理论的关键问题。为在儿童-照护者对话的语法正确性研究中实现更快、更具可重复性且更大规模的语料库分析,自动标注工具可替代繁琐的人工标注。我们提出一种针对儿童-照护者对话中语境依存语法正确性的编码方案,并对大规模转录对话语料库中4000余条话语进行标注。基于这些标注,我们训练并评估了多种自然语言处理模型。结果表明,基于Transformer的微调模型表现最佳,达到人类标注者间的一致性水平。作为该工具的首次应用与合理性检验,我们利用训练后的模型对规模较人工标注数据大近两个数量级的语料库进行标注,验证了儿童语法正确性随年龄增长稳步提升。本研究为运用前沿自然语言处理方法大规模研究儿童语言习得的文献体系做出贡献。