Learning neural subset selection tasks, such as compound selection in AI-aided drug discovery, have become increasingly pivotal across diverse applications. The existing methodologies in the field primarily concentrate on constructing models that capture the relationship between utility function values and subsets within their respective supersets. However, these approaches tend to overlook the valuable information contained within the superset when utilizing neural networks to model set functions. In this work, we address this oversight by adopting a probabilistic perspective. Our theoretical findings demonstrate that when the target value is conditioned on both the input set and subset, it is essential to incorporate an \textit{invariant sufficient statistic} of the superset into the subset of interest for effective learning. This ensures that the output value remains invariant to permutations of the subset and its corresponding superset, enabling identification of the specific superset from which the subset originated. Motivated by these insights, we propose a simple yet effective information aggregation module designed to merge the representations of subsets and supersets from a permutation invariance perspective. Comprehensive empirical evaluations across diverse tasks and datasets validate the enhanced efficacy of our approach over conventional methods, underscoring the practicality and potency of our proposed strategies in real-world contexts.
翻译:学习神经子集选择任务(例如AI辅助药物发现中的化合物选择)已在各类应用中变得日益关键。该领域的现有方法主要集中于构建建模效用函数值与各自超集内子集之间关系的模型。然而,这些方法在利用神经网络建模集合函数时,往往忽略了超集中所包含的宝贵信息。在本工作中,我们通过采用概率视角来解决这一疏漏。我们的理论发现表明,当目标值同时以输入集合和子集为条件时,必须将超集的一个\textit{不变充分统计量}纳入感兴趣的子集,以实现有效学习。这确保了输出值对子集及其对应超集的排列保持不变性,从而能够识别子集源自的具体超集。受这些见解的启发,我们提出一个简单而有效的信息聚合模块,旨在从排列不变性视角融合子集和超集的表示。跨不同任务和数据集的全面实证评估验证了我们的方法相较于传统方法的增强效能,突显了我们提出的策略在真实世界场景中的实用性和效力。