In traditional studies on language evolution, scholars often emphasize the importance of sound laws and sound correspondences for phylogenetic inference of language family trees. However, to date, computational approaches have typically not taken this potential into account. Most computational studies still rely on lexical cognates as major data source for phylogenetic reconstruction in linguistics, although there do exist a few studies in which authors praise the benefits of comparing words at the level of sound sequences. Building on (a) ten diverse datasets from different language families, and (b) state-of-the-art methods for automated cognate and sound correspondence detection, we test, for the first time, the performance of sound-based versus cognate-based approaches to phylogenetic reconstruction. Our results show that phylogenies reconstructed from lexical cognates are topologically closer, by approximately one third with respect to the generalized quartet distance on average, to the gold standard phylogenies than phylogenies reconstructed from sound correspondences.
翻译:在关于语言演化的传统研究中,学者常强调音韵规律与音韵对应对于语言谱系树系统发育推断的重要性。然而迄今为止,计算方法尚未充分考虑这一潜力。大多数计算研究仍依赖词汇同源词作为语言系统发育重建的主要数据源,尽管确有少数研究作者称赞在音韵序列层面比较词汇的优势。基于(一)来自不同语系的十个多样化数据集,以及(二)用于自动识别同源词和音韵对应的最先进方法,我们首次系统检验了基于音韵与基于同源词的系统发育重建方法性能。结果表明:与通过音韵对应重建的系统发育树相比,基于词汇同源词重建的谱系树在拓扑结构上更接近黄金标准系统发育树——通过广义四分体距离平均衡量,其差异程度约缩小三分之一。