Factor analysis (FA) is a statistical tool for studying how observed variables with some mutual dependences can be expressed as functions of mutually independent unobserved factors, and it is widely applied throughout the psychological, biological, and physical sciences. We revisit this classic method from the comparatively new perspective given by advancements in causal discovery and deep learning, introducing a framework for Neuro-Causal Factor Analysis (NCFA). Our approach is fully nonparametric: it identifies factors via latent causal discovery methods and then uses a variational autoencoder (VAE) that is constrained to abide by the Markov factorization of the distribution with respect to the learned graph. We evaluate NCFA on real and synthetic data sets, finding that it performs comparably to standard VAEs on data reconstruction tasks but with the advantages of sparser architecture, lower model complexity, and causal interpretability. Unlike traditional FA methods, our proposed NCFA method allows learning and reasoning about the latent factors underlying observed data from a justifiably causal perspective, even when the relations between factors and measurements are highly nonlinear.
翻译:因子分析(FA)是一种统计工具,用于研究具有相互依赖性的观测变量如何表达为相互独立的未观测因子的函数,广泛应用于心理学、生物学和物理科学领域。我们从因果发现与深度学习进展所提供的新视角重新审视这一经典方法,提出了一套神经因果因子分析(NCFA)框架。我们的方法完全非参数化:通过潜在因果发现方法识别因子,然后使用变分自编码器(VAE),并约束其遵循相对于所学习图结构的分布马尔可夫分解。我们在真实与合成数据集上评估了NCFA,发现其在数据重构任务中性能与标准VAE相当,但具有更稀疏的结构、更低的模型复杂度以及因果可解释性优势。不同于传统FA方法,我们提出的NCFA方法允许从可辩护的因果视角学习和推理观测数据背后的潜在因子,即使因子与测量之间的关系高度非线性。