We introduce and investigate, for finite groups $G$, $G$-invariant deep neural network ($G$-DNN) architectures with ReLU activation that are densely connected-- i.e., include all possible skip connections. In contrast to other $G$-invariant architectures in the literature, the preactivations of the$G$-DNNs presented here are able to transform by \emph{signed} permutation representations (signed perm-reps) of $G$. Moreover, the individual layers of the $G$-DNNs are not required to be $G$-equivariant; instead, the preactivations are constrained to be $G$-equivariant functions of the network input in a way that couples weights across all layers. The result is a richer family of $G$-invariant architectures never seen previously. We derive an efficient implementation of $G$-DNNs after a reparameterization of weights, as well as necessary and sufficient conditions for an architecture to be ``admissible''-- i.e., nondegenerate and inequivalent to smaller architectures. We include code that allows a user to build a $G$-DNN interactively layer-by-layer, with the final architecture guaranteed to be admissible. We show that there are far more admissible $G$-DNN architectures than those accessible with the ``concatenated ReLU'' activation function from the literature. Finally, we apply $G$-DNNs to two example problems -- (1) multiplication in $\{-1, 1\}$ (with theoretical guarantees) and (2) 3D object classification -- % finding that the inclusion of signed perm-reps significantly boosts predictive performance compared to baselines with only ordinary (i.e., unsigned) perm-reps.
翻译:我们针对有限群$G$,引入并研究了采用ReLU激活函数且具有稠密连接(即包含所有可能的跳跃连接)的$G$-不变深度神经网络($G$-DNN)架构。与文献中其他$G$-不变架构不同,本文提出的$G$-DNN的预激活值能够通过$G$的*符号*置换表示(signed perm-reps)进行变换。此外,$G$-DNN的各个层无需满足$G$-等变性;相反,预激活值被约束为网络输入的$G$-等变函数,这种约束以跨所有层耦合权重的方式实现。由此产生了一个前所未有的更丰富的$G$-不变架构族。我们推导出权重重参数化后$G$-DNN的高效实现方法,以及架构“可容许性”(即非退化且不等价于更小架构)的充要条件。我们提供了代码,允许用户逐层交互式构建$G$-DNN,且最终架构保证可容许。我们证明,可容许的$G$-DNN架构数量远超文献中使用“拼接ReLU”激活函数所能实现的架构。最后,我们将$G$-DNN应用于两个示例问题:(1)$\{-1, 1\}$上的乘法(具有理论保证)和(2)三维物体分类——结果发现,与仅使用普通(即无符号)置换表示的基线方法相比,引入符号置换表示显著提升了预测性能。