Machine-Learned Likelihoods (MLL) combines machine-learning classification techniques with likelihood-based inference tests to estimate the experimental sensitivity of high-dimensional data sets. We extend the MLL method by including Kernel Density Estimators (KDE) to avoid binning the classifier output to extract the resulting one-dimensional signal and background probability density functions. We first test our method on toy models generated with multivariate Gaussian distributions, where the true probability distribution functions are known. Later, we apply the method to two cases of interest at the LHC: a search for exotic Higgs bosons, and a $Z'$ boson decaying into lepton pairs. In contrast to physical-based quantities, the typical fluctuations of the ML outputs give non-smooth probability distributions for pure-signal and pure-background samples. The non-smoothness is propagated into the density estimation due to the good performance and flexibility of the KDE method. We study its impact on the final significance computation, and we compare the results using the average of several independent ML output realizations, which allows us to obtain smoother distributions. We conclude that the significance estimation turns out to be not sensible to this issue.
翻译:机器学习似然法(MLL)将机器学习分类技术与基于似然的推断检验相结合,用于评估高维数据集的实验灵敏度。我们通过引入核密度估计(KDE)扩展了MLL方法,以避免对分类器输出进行分箱,从而提取一维信号与背景概率密度函数。首先,我们在已知真实概率分布函数的多变量高斯分布玩具模型上测试该方法。随后,将方法应用于大型强子对撞机(LHC)的两个典型案例:搜寻奇异希格斯玻色子,以及衰变至轻子对的$Z'$玻色子。与基于物理量的结果不同,机器学习输出的典型波动会导致纯信号与纯背景样本的概率分布不光滑。由于KDE方法性能优越且灵活,这种非光滑性会传递至密度估计中。我们研究了该问题对最终显著性计算的影响,并通过比较多次独立机器学习输出实现的平均结果,验证了获得更光滑分布的可能性。最终结论表明,显著性估计对该问题不敏感。