The standard Gaussian Process (GP) only considers a single output sample per input in the training set. Datasets for subjective tasks, such as spoken language assessment, may be annotated with output labels from multiple human raters per input. This paper proposes to generalise the GP to allow for these multiple output samples in the training set, and thus make use of available output uncertainty information. This differs from a multi-output GP, as all output samples are from the same task here. The output density function is formulated to be the joint likelihood of observing all output samples, and latent variables are not repeated to reduce computation cost. The test set predictions are inferred similarly to a standard GP, with a difference being in the optimised hyper-parameters. This is evaluated on speechocean762, showing that it allows the GP to compute a test set output distribution that is more similar to the collection of reference outputs from the multiple human raters.
翻译:标准高斯过程(GP)在训练集中仅考虑每个输入对应一个输出样本。对于主观任务(如口语语言评估)的数据集,每个输入可能由多位人工评分者标注输出标签。本文提出对高斯过程进行推广,使其能够处理训练集中这些多个输出样本,从而利用可用的输出不确定性信息。这与多输出高斯过程不同,因为此处所有输出样本均来自同一任务。输出密度函数被构建为观测所有输出样本的联合似然,且为避免计算成本增加,潜变量不重复设置。测试集预测的推理方式与标准高斯过程类似,区别在于优化后的超参数。该方法在speechocean762数据集上进行评估,结果表明它能使高斯过程计算出的测试集输出分布更接近多位人工评分者提供的参考输出集合。