The unstructured nature of data used in foundation model development is a challenge to systematic analyses for making data use and documentation decisions. From a Responsible AI perspective, these decisions often rely upon understanding how people are represented in data. We propose a framework designed to guide analysis of human representation in unstructured data and identify downstream risks. We apply the framework in two toy examples using the Common Crawl web text corpus (C4) and LAION-400M. We also propose a set of hypothetical action steps in service of dataset use, development, and documentation.
翻译:基础模型开发所用数据的非结构化特性,对开展数据使用与文档化决策的系统性分析构成了挑战。从负责任人工智能的视角出发,这类决策往往依赖于理解数据中人群的表征方式。我们提出一个旨在指导非结构化数据中人群表征分析的框架,并识别下游风险。我们以Common Crawl网络文本语料库(C4)和LAION-400M为两个示例,对该框架进行了应用。同时,我们提出一套服务于数据集使用、开发与文档化的假设性行动步骤。