Hallucination detection in large language and vision-language models is increasingly framed as selective prediction, where a detector assigns a confidence score and abstains when confidence is low. Unsupervised sampling detectors (Semantic Entropy) avoid labels but plateau in quality, while supervised probes attain stronger in-distribution scores yet degrade sharply when calibration labels are scarce. We recover the response manifold of an LLM as the density ridge of a kernel density estimate built on a six-dimensional kinematic feature map of hidden state generation trajectories. A test generation is scored by the negated Euclidean distance from its projected feature point to the nearest ridge vertex, yielding a low-dimensional geometric skeleton of the stochastic output distribution. We evaluate against Semantic Entropy, topological methods, and log-probability on six QA benchmarks (HaluEval-QA, TriviaQA, GSM8K, POPE, ScienceQA, A-OKVQA) using eight text and vision LLMs in a deliberately label-scarce protocol ($n_{\text{cal}}{=}200$ queries, $N{=}5$ generations). Our ridge-based score beats on AUROC with 5-20 points gain, while demonstrating tempered degradation under calibration-label scarcity.
翻译:大语言模型与视觉语言模型中的幻觉检测日益被建模为选择性预测问题,即检测器在置信度较低时分配置信度分数并进行弃权。无监督采样检测器(语义熵)虽避免使用标签但质量陷入平台期,而有监督探针虽在分布内取得更优分数,但在校准标签稀疏时性能急剧下降。我们将大语言模型的响应流形重构为基于隐藏状态生成轨迹六维运动学特征图的核密度估计的密度脊。通过将测试生成投影特征点到最近脊顶点的负欧氏距离进行评分,得到随机输出分布的低维几何骨架。我们在刻意构造的标签稀疏协议(校准查询数=200次、生成数=5次)下,使用八个文本与视觉大语言模型对六个问答基准(HaluEval-QA、TriviaQA、GSM8K、POPE、ScienceQA、A-OKVQA)进行评估,并与语义熵、拓扑方法及对数概率方法进行比较。基于脊的评分在AUROC指标上取得5-20个百分点的提升,同时在校准标签稀疏条件下表现出缓和的性能退化。