Recent work has formulated the task for computational construction grammar as producing a constructicon given a corpus of usage. Previous work has evaluated these unsupervised grammars using both internal metrics (for example, Minimum Description Length) and external metrics (for example, performance on a dialectology task). This paper instead takes a linguistic approach to evaluation, first learning a constructicon and then analyzing its contents from a linguistic perspective. This analysis shows that a learned constructicon can be divided into nine major types of constructions, of which Verbal and Nominal are the most common. The paper also shows that both the token and type frequency of constructions can be used to model variation across registers and dialects.
翻译:近期研究将计算构式语法的任务表述为:基于用法语料库生成构建库。先前的工作通过内部指标(例如最小描述长度)和外部指标(例如方言学任务中的表现)评估这些无监督语法系统。本文则采取语言学视角进行评估,首先学习一个构建库,再从语言学角度分析其内容。分析表明,习得的构建库可划分为九种主要构式类型,其中动词构式和名词构式最为常见。本文还论证了构式的词例频率与类型频率均可用于建模不同语域和方言间的差异。