Machine learning tools often rely on embedding text as vectors of real numbers. In this paper, we study how the semantic structure of language is encoded in the algebraic structure of such embeddings. Specifically, we look at a notion of ``semantic independence'' capturing the idea that, e.g., ``eggplant'' and ``tomato'' are independent given ``vegetable''. Although such examples are intuitive, it is difficult to formalize such a notion of semantic independence. The key observation here is that any sensible formalization should obey a set of so-called independence axioms, and thus any algebraic encoding of this structure should also obey these axioms. This leads us naturally to use partial orthogonality as the relevant algebraic structure. We develop theory and methods that allow us to demonstrate that partial orthogonality does indeed capture semantic independence. Complementary to this, we also introduce the concept of independence preserving embeddings where embeddings preserve the conditional independence structures of a distribution, and we prove the existence of such embeddings and approximations to them.
翻译:机器学习工具通常将文本表示为实数向量。本文研究语言的语义结构如何编码在嵌入的代数结构中。具体而言,我们探讨"语义独立性"这一概念,该概念捕捉了诸如"茄子"和"番茄"在给定"蔬菜"条件下相互独立的思想。尽管此类示例直观易懂,但形式化这种语义独立性概念却具有挑战性。关键发现是:任何合理的形式化都必须服从一组所谓的独立性公理,因此对这种结构的任何代数编码也应遵循这些公理。这自然引导我们采用部分正交性作为相关的代数结构。我们开发了相应的理论与方法,证明部分正交性确实能够捕捉语义独立性。作为补充,我们还引入了"独立性保持嵌入"的概念——即嵌入能够保留分布的条件独立结构,并证明了此类嵌入及其近似形式的存在性。