Human annotator simulation (HAS) serves as a cost-effective substitute for human evaluation such as data annotation and system assessment. Human perception and behaviour during human evaluation exhibit inherent variability due to diverse cognitive processes and subjective interpretations, which should be taken into account in modelling to better mimic the way people perceive and interact with the world. This paper introduces a novel meta-learning framework that treats HAS as a zero-shot density estimation problem, which incorporates human variability and allows for the efficient generation of human-like annotations for unlabelled test inputs. Under this framework, we propose two new model classes, conditional integer flows and conditional softmax flows, to account for ordinal and categorical annotations, respectively. The proposed method is evaluated on three real-world human evaluation tasks and shows superior capability and efficiency to predict the aggregated behaviours of human annotators, match the distribution of human annotations, and simulate the inter-annotator disagreements.
翻译:人类标注者模拟(HAS)作为人类评估(如数据标注和系统评估)的经济高效替代方案。在人类评估过程中,由于认知过程差异和主观解释,人类感知与行为表现出固有变异性,建模时应考虑这些因素以更好地模仿人类感知和与世界互动的方式。本文提出一种新颖的元学习框架,将HAS视为零样本密度估计问题,该框架融合人类变异性,并允许为未标记测试输入高效生成类人标注。在此框架下,我们提出两类新模型:条件整数流和条件softmax流,分别处理有序标注和分类标注。该方法在三个真实世界人类评估任务上进行了评估,展现出预测人类标注者聚合行为、匹配人类标注分布以及模拟标注者间分歧的卓越能力和效率。