Can we identify the parameters of a neural network by probing its input-output mapping? Usually, there is no unique solution because of permutation, overparameterisation and activation function symmetries. Yet, we show that the incoming weight vector of each neuron is identifiable up to sign or scaling, depending on the activation function. For all commonly used activation functions, our novel method 'Expand-and-Cluster' identifies the size and parameters of a target network in two phases: (i) to relax the non-convexity of the problem, we train multiple student networks of expanded size to imitate the mapping of the target network; (ii) to identify the target network, we employ a clustering procedure and uncover the weight vectors shared between students. We demonstrate successful parameter and size recovery of trained shallow and deep networks with less than 10% overhead in the neuron number and describe an 'ease-of-identifiability' axis by analysing 150 synthetic problems of variable difficulty.
翻译:能否通过探测神经网络的输入-输出映射来识别其参数?通常,由于排列、过参数化和激活函数对称性的存在,解并不唯一。然而,我们证明每个神经元的输入权重向量在符号或尺度上是可识别的,具体取决于激活函数。针对所有常用激活函数,我们提出的新方法“扩展与聚类”通过两个阶段来识别目标网络的规模和参数:(i)为放松问题的非凸性,我们训练多个扩展规模的学生网络来模仿目标网络的映射;(ii)为识别目标网络,我们采用聚类过程,揭示学生网络之间共享的权重向量。我们成功实现了对训练过的浅层和深层网络的参数与规模恢复,神经元数量开销低于10%,并通过分析150个难度各异的人工合成问题,描述了“可识别性难易度”轴。