The ubiquity of Big Data and machine learning in society evinces the need of further investigation of their fundamental limitations. In this paper, we extend the ``too-much-information-tends-to-behave-like-very-little-information'' phenomenon to formal knowledge about lawlike universes and arbitrary collections of computably generated datasets. This gives rise to the simplicity bubble problem, which refers to a learning algorithm equipped with a formal theory that can be deceived by a dataset to find a locally optimal model which it deems to be the global one. However, the actual high-complexity globally optimal model unpredictably diverges from the found low-complexity local optimum. Zemblanity is defined by an undesirable but expected finding that reveals an underlying problem or negative consequence in a given model or theory, which is in principle predictable in case the formal theory contains sufficient information. Therefore, we argue that there is a ceiling above which formal knowledge cannot further decrease the probability of zemblanitous findings, should the randomly generated data made available to the learning algorithm and formal theory be sufficiently large in comparison to their joint complexity.
翻译:大数据与机器学习在社会中的普遍存在,凸显了进一步探究其根本局限性的必要性。本文将“信息过多往往表现得如同信息过少”这一现象,拓展至关于类律宇宙的形式化知识及可计算生成数据集的任意集合。由此引出简单性泡沫问题:当学习算法配备形式化理论时,可能被数据集欺骗,从而找到其视为全局最优的局部最优模型;然而,实际的高复杂度全局最优模型却不可预测地偏离了所发现的低复杂度局部最优解。泽姆布拉尼特性被定义为一种不期望但可预期的发现,它揭示了给定模型或理论中潜在的问题或负面后果——原则上,若形式化理论包含足够信息,该发现是可预测的。因此,我们论证:若随机生成的数据集相对于学习算法与形式化理论的联合复杂度足够大,则形式化知识存在一个上限,超越该上限便无法进一步降低出现泽姆布拉尼特发现的可能性。