Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate uncertainty in black-box settings, yet the confidence scale itself (typically 0--100) is rarely examined. We show that this design choice is not neutral. Across six LLMs and three datasets, verbalized confidence is heavily discretized, with more than 78\% of responses concentrating on just three round-number values. To investigate this phenomenon, we systematically manipulate confidence scales along three dimensions: granularity, boundary placement, and range regularity, and evaluate metacognitive sensitivity using $meta\text{-}d'$. We find that a 0--20 scale consistently improves metacognitive efficiency over the standard 0--100 format, while boundary compression degrades performance and round-number preferences persist even under irregular ranges. These results demonstrate that confidence scale design directly affects the quality of verbalized uncertainty and should be treated as a first-class experimental variable in LLM evaluation.
翻译:口头化置信度(即大语言模型报告数值化确定性分数)被广泛用于估计黑箱环境中的不确定性,然而置信度尺度本身(通常为0-100)却鲜少受到审视。我们证明这一设计选择并非中性。在六个大语言模型和三个数据集上,口头化置信度呈现高度离散化特征——超过78%的响应集中在仅三个整数数值上。为探究此现象,我们沿三个维度系统操控置信度尺度:粒度、边界位置及范围规则性,并使用$meta\text{-}d'$评估元认知敏感性。研究发现:0-20尺度相较标准0-100格式能持续提升元认知效率,边界压缩会降低性能,而即使在非规则范围下整数偏好依然存在。这些结果表明,置信度尺度设计直接影响口头化不确定性的质量,应作为大语言模型评估中的首要实验变量。