Guidance on how to validate computational text-based measures of social science constructs is fragmented. Whereas scholars are generally acknowledging the importance of validating their text-based measures, they often lack common terminology and a unified framework to do so. This paper introduces a new validation framework called ValiTex, designed to assist scholars to measure social science constructs based on textual data. The framework draws on a long-established tradition within psychometrics while extending the framework for the purpose of computational text analysis. ValiTex consists of two components, a conceptual model, and a dynamic checklist. Whereas the conceptual model provides a general structure along distinct phases on how to approach validation, the dynamic checklist defines specific validation steps and provides guidance on which steps might be considered recommendable (i.e., providing relevant and necessary validation evidence) or optional (i.e., useful for providing additional supporting validation evidence. The utility of the framework is demonstrated by applying it to a use case of detecting sexism from social media data.
翻译:关于如何验证社会科学构念基于文本的计算测量,现有指导较为零散。尽管学者们普遍承认验证文本测量方法的重要性,但他们往往缺乏通用术语和统一框架来开展此项工作。本文提出一种名为ValiTex的新验证框架,旨在帮助学者基于文本数据测量社会科学构念。该框架借鉴心理测量学中历史悠久的传统,同时针对计算文本分析的目标对其进行了扩展。ValiTex包含两个组成部分:概念模型与动态检查表。概念模型围绕不同阶段提供了如何开展验证的总体结构,而动态检查表则定义了具体的验证步骤,并指导哪些步骤应被视为推荐项(即提供相关且必要的验证证据)或可选项目(即有助于提供额外的辅助验证证据)。通过将该框架应用于从社交媒体数据中检测性别歧视的案例,展示了其实用价值。