Adversarial vulnerability in vision and hallucination in large language models are conventionally viewed as separate problems, each addressed with modality-specific patches. This study first reveals that they share a common geometric origin: the input and its loss gradient are conjugate observables subject to an irreducible uncertainty bound. Formalizing a Neural Uncertainty Principle (NUP) under a loss-induced state, we find that in near-bound regimes, further compression must be accompanied by increased sensitivity dispersion (adversarial fragility), while weak prompt-gradient coupling leaves generation under-constrained (hallucination). Crucially, this bound is modulated by an input-gradient correlation channel, captured by a specifically designed single-backward probe. In vision, masking highly coupled components improves robustness without costly adversarial training; in language, the same prefill-stage probe detects hallucination risk before generating any answer tokens. NUP thus turns two seemingly separate failure taxonomies into a shared uncertainty-budget view and provides a principled lens for reliability analysis. Guided by this NUP theory, we propose ConjMask (masking high-contribution input components) and LogitReg (logit-side regularization) to improve robustness without adversarial training, and use the probe as a decoding-free risk signal for LLMs, enabling hallucination detection and prompt selection. NUP thus provides a unified, practical framework for diagnosing and mitigating boundary anomalies across perception and generation tasks.
翻译:视觉中的对抗脆弱性和大语言模型中的幻觉常规上被视为独立问题,各自依赖模态特定补丁加以解决。本研究首先揭示二者共享共同的几何起源:输入及其损失梯度是共轭可观测量,受制于不可约的不确定度下界。在损失诱导态下形式化神经不确定性原理(Neural Uncertainty Principle, NUP)后,我们发现:在近边界区域内,进一步压缩必然伴随敏感性离散度增大(即对抗脆弱性),而弱提示-梯度耦合则导致生成过程欠约束(即幻觉)。关键的是,该下界受输入-梯度相关性信道调制,可通过专门设计的单反向传播探针捕获。在视觉领域,掩蔽高耦合成分能提升鲁棒性,无需昂贵对抗训练;在语言领域,同一预填充阶段探针可在生成任何答案词元前检测幻觉风险。因此,NUP将两种看似独立的故障分类转化为统一的不确定度预算视角,并为可靠性分析提供原理性透镜。在NUP理论指导下,我们提出ConjMask(掩蔽高贡献输入成分)和LogitReg(logit侧正则化)以提升鲁棒性,免于对抗训练,并将该探针用作大语言模型的无解码风险信号,实现幻觉检测与提示选择。由此,NUP为跨感知与生成任务的边界异常诊断与缓解提供了统一且实用的框架。