Evidential deep learning, built upon belief theory and subjective logic, offers a principled and computationally efficient way to turn a deterministic neural network uncertainty-aware. The resultant evidential models can quantify fine-grained uncertainty using the learned evidence. To ensure theoretically sound evidential models, the evidence needs to be non-negative, which requires special activation functions for model training and inference. This constraint often leads to inferior predictive performance compared to standard softmax models, making it challenging to extend them to many large-scale datasets. To unveil the real cause of this undesired behavior, we theoretically investigate evidential models and identify a fundamental limitation that explains the inferior performance: existing evidential activation functions create zero evidence regions, which prevent the model to learn from training samples falling into such regions. A deeper analysis of evidential activation functions based on our theoretical underpinning inspires the design of a novel regularizer that effectively alleviates this fundamental limitation. Extensive experiments over many challenging real-world datasets and settings confirm our theoretical findings and demonstrate the effectiveness of our proposed approach.
翻译:基于信念理论和主观逻辑的深度证据学习,为将确定性神经网络转化为不确定性感知模型提供了一种原理性且计算高效的方法。由此产生的证据模型能够利用学习到的证据量化细粒度不确定性。为确保证据模型在理论上成立,证据需满足非负性要求,这需要训练和推理过程中采用特殊激活函数。然而该约束常导致模型预测性能低于标准Softmax模型,使其难以拓展至大规模数据集。为揭示这一不良行为的真正成因,我们对证据模型展开理论分析,并发现一个解释其性能劣势的根本性局限:现有证据激活函数会产生零证据区域,阻碍模型从落入该区域的训练样本中学习。基于理论基础的深度分析启发我们设计了一种新型正则化器,可有效缓解这一根本性局限。在多个具有挑战性的真实数据集和场景下进行的广泛实验,既验证了理论发现,也证明了所提方法的有效性。