In Radhakrishnan et al. [2020], the authors empirically show that autoencoders trained with usual SGD methods shape out basins of attraction around their training data. We consider network functions of width not exceeding the input dimension and prove that in this situation basins of attraction are bounded and their complement cannot have bounded components. Our conditions in these results are met in several experiments of the latter work and we thus address a question posed therein. We also show that under some more restrictive conditions the basins of attraction are path-connected. The tightness of the conditions in our results is demonstrated by means of several examples. Finally, the arguments used to prove the above results allow us to derive a root cause why scalar-valued neural network functions that fulfill our bounded width condition are not dense in spaces of continuous functions.
翻译:在Radhakrishnan等人[2020]的研究中,作者通过实验表明,采用常规随机梯度下降方法训练的自动编码器会在训练数据周围形成吸引盆。我们考虑宽度不超过输入维度的网络函数,并证明在此情形下吸引盆是有界的,且其补集不存在有界连通分量。我们结论中的条件在该工作的多项实验中均得到满足,从而回应了其中提出的一个疑问。我们还证明,在更严格的条件下,这些吸引盆是道路连通的。通过多个实例展示了我们结论中条件的紧致性。最后,用于证明上述结论的论证方法使我们能够推导出满足有界宽度条件的标量值神经网络函数在连续函数空间中非稠密的根本原因。