We propose \textit{masked particle modeling} (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.
翻译:我们提出掩码粒子建模(MPM)作为一种自监督方法,用于学习无序输入集合上的通用、可迁移且可复用的表征,以应用于高能物理(HEP)科学数据。本研究提出了一种基于掩码建模的预训练新方案,以学习集合上的置换不变函数。更广泛地,这项工作为构建HEP大型基础模型迈出了重要一步,此类模型可通过自监督学习进行通用预训练,随后针对各类下游任务进行微调。在MPM中,集合中的粒子被掩码,训练目标是通过预训练向量量化变分自编码器生成的离散化令牌表征恢复其身份。我们通过对撞机物理实验中高能喷注样本验证该方法的有效性,包括离散化、置换不变性及排序影响的研究。我们还研究了模型的微调能力,证明其可适应监督与弱监督喷注分类等任务,且模型能够通过少量微调数据集高效迁移至新类别与新数据领域。