As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks face the Construct-Composition-Context ($C^3$) challenge: relying on discriminative, multiple-choice formats that probe value knowledge rather than true orientations, overlook subcultural heterogeneity, and mismatch with real-world open-ended generation. We introduce DOVE, a distributional evaluation framework that directly compares human-written text distributions with LLM-generated outputs. DOVE utilizes a rate-distortion variational optimization objective to construct a compact value codebook from 10K documents, mapping text into a structured value space to filter semantic noise. Alignment is measured using unbalanced optimal transport, capturing intra-cultural distributional structures and subgroup diversity. Experiments across 12 LLMs show that DOVE achieves superior predictive validity, attaining a 31.56% correlation with downstream tasks, while maintaining high reliability with as few as 500 samples per culture.
翻译:随着大语言模型的全球部署,对齐其文化价值取向对安全性和用户参与度至关重要。然而,现有基准面临构造-组成-上下文(C³)挑战:依赖判别性多选题格式,仅探测价值知识而非真实取向,忽视子文化异质性,并偏离现实世界的开放式生成场景。我们提出DOVE(分布式评估框架),该框架直接比较人类撰写的文本分布与大语言模型生成的输出。DOVE利用率失真变分优化目标,从10K文档中构建紧凑的价值码本,将文本映射到结构化的价值空间以滤除语义噪声。通过非平衡最优传输测量对齐度,捕捉文化内分布结构与子群体多样性。在12个大语言模型上的实验表明,DOVE在仅需每文化500个样本的条件下,下游任务相关度达31.56%且具有优异预测效度,同时保持高可靠性。