We explore the possibility of meta-learning for the language-independent unsupervised tokenization problem for English, Russian, and Chinese. We implement the meta-learning approach for automatic determination of hyper-parameters of the unsupervised tokenization model proposed in earlier works, relying on various human-independent fitness functions such as normalised anti-entropy, compression factor and cross-split F1 score, as well as additive and multiplicative composite combinations of the three metrics, testing them against the conventional F1 tokenization score. We find a fairly good correlation between the latter and the additive combination of the former three metrics for English and Russian. In case of Chinese, we find a significant correlation between the F 1 score and the compression factor. Our results suggest the possibility of robust unsupervised tokenization of low-resource and dead languages and allow us to think about human languages in terms of the evolution of efficient symbolic communication codes with different structural optimisation schemes that have evolved in different human cultures.
翻译:我们探索了元学习在英语、俄语和汉语的无语言依赖无监督分词问题中的应用。我们实现了元学习方法来自动确定早期研究中提出的无监督分词模型的超参数,该方法依赖于多种不依赖人工的适应度函数,如归一化反熵、压缩因子和交叉分割F1分数,以及这三种指标的加性和乘性复合组合,并将它们与传统F1分词分数进行比较。我们发现,对于英语和俄语,传统F1分数与前三项指标的加性组合之间存在相当好的相关性。对于汉语,我们发现F1分数与压缩因子之间存在显著相关性。我们的结果表明,对低资源语言和死语言进行稳健的无监督分词是可能的,并使我们能够从高效符号通信代码(这些代码经过不同人类文化中产生的不同结构优化方案演化而来)的角度思考人类语言。