Conformal prediction has shown spurring performance in constructing statistically rigorous prediction sets for arbitrary black-box machine learning models, assuming the data is exchangeable. However, even small adversarial perturbations during the inference can violate the exchangeability assumption, challenge the coverage guarantees, and result in a subsequent decline in empirical coverage. In this work, we propose a certifiably robust learning-reasoning conformal prediction framework (COLEP) via probabilistic circuits, which comprise a data-driven learning component that trains statistical models to learn different semantic concepts, and a reasoning component that encodes knowledge and characterizes the relationships among the trained models for logic reasoning. To achieve exact and efficient reasoning, we employ probabilistic circuits (PCs) within the reasoning component. Theoretically, we provide end-to-end certification of prediction coverage for COLEP in the presence of bounded adversarial perturbations. We also provide certified coverage considering the finite size of the calibration set. Furthermore, we prove that COLEP achieves higher prediction coverage and accuracy over a single model as long as the utilities of knowledge models are non-trivial. Empirically, we show the validity and tightness of our certified coverage, demonstrating the robust conformal prediction of COLEP on various datasets, including GTSRB, CIFAR10, and AwA2. We show that COLEP achieves up to 12% improvement in certified coverage on GTSRB, 9% on CIFAR-10, and 14% on AwA2.
翻译:共形预测在假设数据可交换的前提下,为任意黑箱机器学习模型构建统计严谨的预测集方面展现了显著性能。然而,推理过程中的微小对抗扰动可能破坏可交换性假设,挑战覆盖保证,并导致经验覆盖率的下降。本文提出一种基于概率电路的认证鲁棒学习-推理共形预测框架(COLEP),其包含数据驱动的学习组件(训练统计模型以学习不同语义概念)和推理组件(编码知识并表征已训练模型间关系以进行逻辑推理)。为实现精确高效的推理,我们在推理组件中采用概率电路(PC)。理论上,我们针对有界对抗扰动下的COLEP预测覆盖提供了端到端认证,并考虑了校准集有限规模下的认证覆盖。此外,我们证明只要知识模型的效用非平凡,COLEP相比单一模型能实现更高的预测覆盖率和准确度。实验层面,我们展示了认证覆盖的有效性与紧致性,并在GTSRB、CIFAR10和AwA2等数据集上验证了COLEP的鲁棒共形预测能力。结果表明,COLEP在GTSRB、CIFAR-10和AwA2上分别实现了高达12%、9%和14%的认证覆盖率提升。