Aligning model representations to humans has been found to improve robustness and generalization. However, such methods often focus on standard observational data. Synthetic data is proliferating and powering many advances in machine learning; yet, it is not always clear whether synthetic labels are perceptually aligned to humans -- rendering it likely model representations are not human aligned. We focus on the synthetic data used in mixup: a powerful regularizer shown to improve model robustness, generalization, and calibration. We design a comprehensive series of elicitation interfaces, which we release as HILL MixE Suite, and recruit 159 participants to provide perceptual judgments along with their uncertainties, over mixup examples. We find that human perceptions do not consistently align with the labels traditionally used for synthetic points, and begin to demonstrate the applicability of these findings to potentially increase the reliability of downstream models, particularly when incorporating human uncertainty. We release all elicited judgments in a new data hub we call H-Mix.
翻译:研究发现,使模型表征与人类对齐能够提升模型的鲁棒性和泛化能力。然而,这类方法通常聚焦于标准观测数据。合成数据正蓬勃发展并推动着机器学习领域的诸多进步,但合成标签是否在感知层面与人类对齐尚不明确——这可能导致模型表征无法实现人类对齐。本研究聚焦于混合增强(mixup)中使用的合成数据:该数据是一种已被证实能提升模型鲁棒性、泛化能力和校准性能的强大正则化手段。我们设计了一套全面的感知触发接口(并将其开源为HILL MixE Suite),招募159名参与者针对混合增强样本提供带有不确定性的感知判断。研究发现,人类对合成数据点的感知与传统标签并不一致,并初步证明了这些发现可用于提升下游模型可靠性(尤其在引入人类不确定性因素时)的可行性。我们将所有采集的判断数据发布至名为H-Mix的新数据枢纽。