Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. Answering this for natural language is difficult: tasks are hard to delineate and can confound one another. To rigorously investigate the relationship between data frequency and learnability, we turn to a controlled setting using formal languages induced from probabilistic finite automata. These serve as a methodological testbed to demonstrate that standard correlational evaluation practices are inherently flawed. To enable causal analysis, we introduce the binning semiring, an algebraic object that lets us control how often a targeted property occurs in a sampled corpus. We formulate the experimental pipeline as a causal graphical model and derive decomposed Kullback-Leibler divergence metrics to measure the learnability of specific sub-tasks. Our experiments show that evaluating learnability without causal intervention leads to incorrect conclusions due to confounders in correlational analysis, and serve as a warning about correlational pitfalls in natural-language settings.
翻译:作为多任务学习者的语言模型在训练过程中获取广泛能力,其中一个根本问题是:学习特定任务需要多少任务专属数据?针对自然语言回答该问题较为困难——任务边界难以界定且可能相互混淆。为严格探究数据频率与可学性之间的关系,我们转向基于概率有限自动机诱导的形式语言这一受控场景。该场景作为方法论测试平台,证明了标准关联性评估实践存在固有缺陷。为开展因果分析,我们引入分箱半环这一代数对象,使其能够控制目标属性在采样语料库中出现的频率。我们将实验流程形式化为因果图模型,并推导出分解型KL散度指标以衡量特定子任务的可学性。实验表明:若不进行因果干预而直接评估可学性,会因关联性分析中的混杂因子导致错误结论,这为自然语言场景中的关联性陷阱提供了警示。