Marginal inference in discrete graphical models forces a choice between exactness and scalability: exact algorithms are intractable for high-treewidth graphs, while iterative approximations (Belief Propagation, variational methods) sacrifice convergence guarantees on frustrated topologies. We argue that this dichotomy stems from a mismatched inductive bias: iterative methods abandon the sequential elimination structure that makes exact inference correct. We introduce In-Context Graphical Inference (ICG-I), an autoregressive Graph Transformer that restores this structure by mimicking Variable Elimination with learned, Tensor- Train-compressed intermediate factors, paired with a Dirichlet output layer and Weighted Conformal Prediction for calibrated, distribution-free coverage guarantees under topological shift. We prove that TT compression errors propagate at most lincarly through the autoregressive chain, that the Dirichlet-Multinomial loss is a proper scoring rule, and that WCP maintains coverage with a quantifiable degradation under estimated density ratios. We conducted intensive experiments to evaluate ICG-I and achieved state-of-the-art performance across all benchmarks. ICG-I reduces MAE from 0.041 (best baseline) to 0.020 on standard instances and achieves 0.048 on N=500 frustrated spin glasses where BP diverges entirely.
翻译:离散图形模型中的边缘推理在精确性和可扩展性之间迫使我们做出选择:精确算法对于高树宽图难以处理,而迭代近似方法(信念传播、变分方法)在受阻拓扑上牺牲了收敛保证。我们认为,这种二分法源于不匹配的归纳偏差:迭代方法放弃了使精确推理正确的顺序消元结构。我们提出了上下文图形推理(ICG-I),这是一种自回归图形Transformer,通过模仿变量消元并使用学习到的张量训练压缩中间因子来恢复这种结构,同时结合狄利克雷输出层和加权共形预测,以在拓扑变化下提供校准的、无分布假设的覆盖保证。我们证明,TT压缩误差在自回归链中最多线性传播,狄利克雷-多项损失是一种适当的评分规则,并且WCP在估计密度比下保持覆盖且退化可量化。我们通过大量实验评估了ICG-I,并在所有基准测试中取得了最先进的性能。ICG-I将标准实例上的MAE从0.041(最佳基线)降至0.020,并在N=500的受阻自旋玻璃系统中(BP完全发散)实现了0.048。