As generative AI technologies are increasingly being launched across the globe, assessing their competence to operate in different cultural contexts is exigently becoming a priority. While recent years have seen numerous and much-needed efforts on cultural benchmarking, these efforts have largely focused on specific aspects of culture and evaluation. While these efforts contribute to our understanding of cultural competence, a unified and systematic evaluation approach is needed for us as a field to comprehensively assess diverse cultural dimensions at scale. Drawing on measurement theory, we present a principled framework to aggregate multifaceted indicators of cultural capabilities into a unified assessment of cultural intelligence. We start by developing a working definition of culture that includes identifying core domains of culture. We then introduce a broad-purpose, systematic, and extensible framework for assessing cultural intelligence of AI systems. Drawing on theoretical framing from psychometric measurement validity theory, we decouple the background concept (i.e., cultural intelligence) from its operationalization via measurement. We conceptualize cultural intelligence as a suite of core capabilities spanning diverse domains, which we then operationalize through a set of indicators designed for reliable measurement. Finally, we identify the considerations, challenges, and research pathways to meaningfully measure these indicators, specifically focusing on data collection, probing strategies, and evaluation metrics.
翻译:随着生成式人工智能技术在全球范围内日益推广,评估其在多元文化环境中运作的能力正变得迫在眉睫。尽管近年来在文化基准测试方面开展了大量且十分必要的工作,但这些工作大多聚焦于文化的特定方面和评估。虽然这些努力增进了我们对文化能力的理解,但该领域需要一个统一且系统的评估方法,以便大规模地全面评估不同的文化维度。借鉴测量理论,我们提出了一个基于原则的框架,将文化能力的多面指标整合为对文化智力的统一评估。我们首先对文化进行工作定义,包括识别文化的核心领域。随后,我们引入一个通用、系统且可扩展的框架,用于评估人工智能系统的文化智力。借鉴心理测量效度理论的理论框架,我们将背景概念(即文化智力)与其通过测量的操作化过程相分离。我们将文化智力概念化为涵盖不同领域的核心能力组合,然后通过一套为可靠测量而设计的指标对其进行操作化。最后,我们确定了有意义地测量这些指标时需考虑的要素、面临的挑战以及研究路径,特别关注数据收集、探测策略和评估指标。