While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: the uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.
翻译:虽然自然语言在典型词序和词序灵活性上存在显著差异,但其词序仍遵循共同的跨语言统计模式,这通常归因于功能性压力。在识别这些压力的研究中,已有工作比较了真实词序与反事实词序。然而,此类研究忽视了一个功能性压力:均匀信息密度假说,该假说认为信息应均匀分布在话语中。本研究探究均匀信息密度的压力是否跨语言地影响了词序模式。为此,我们使用计算模型检验真实词序是否比反事实词序带来更高的信息均匀性。通过对10种类型学多样语言的实证研究,我们发现:(i) 在SVO语言中,真实词序的信息均匀性始终优于逆向词序;(ii) 只有语言学上不合理的反事实词序才始终超越真实词序的均匀性。这些结果与自然语言发展使用中存在信息均匀性压力的观点一致。