Neural networks excel across a wide range of tasks, yet remain black boxes. In particular, how their internal representations are shaped by the complexity of the input data and the problems they solve remains obscure. In this work, we introduce a suite of five data-agnostic probes-pruning, binarization, noise injection, sign flipping, and bipartite network randomization-to quantify how task difficulty influences the topology and robustness of representations in multilayer perceptrons (MLPs). MLPs are represented as signed, weighted bipartite graphs from a network science perspective. We contrast easy and hard classification tasks on the MNIST and Fashion-MNIST datasets. We show that binarizing weights in hard-task models collapses accuracy to chance, whereas easy-task models remain robust. We also find that pruning low-magnitude edges in binarized hard-task models reveals a sharp phase-transition in performance. Moreover, moderate noise injection can enhance accuracy, resembling a stochastic-resonance effect linked to optimal sign flips of small-magnitude weights. Finally, preserving only the sign structure-instead of precise weight magnitudes-through bipartite network randomizations suffices to maintain high accuracy. These phenomena define a model- and modality-agnostic measure of task complexity: the performance gap between full-precision and binarized or shuffled neural network performance. Our findings highlight the crucial role of signed bipartite topology in learned representations and suggest practical strategies for model compression and interpretability that align with task complexity.
翻译:神经网络在广泛任务中表现出色,但其内部运作仍属黑箱。尤其,输入数据复杂性及所解决问题如何塑造其内部表征仍不明确。本文引入五项与数据无关的探测方法——剪枝、二值化、噪声注入、符号翻转及二分网络随机化——以量化任务难度如何影响多层感知机(MLP)中表征的拓扑结构与鲁棒性。从网络科学视角,将MLP表示为带符号加权的二分图。我们对比MNIST与Fashion-MNIST数据集上的简单与困难分类任务。研究表明:困难任务模型的权重二值化会使准确率骤降至随机水平,而简单任务模型仍保持鲁棒。进一步发现,对二值化困难任务模型进行低幅值边剪枝会引发性能的尖锐相变。此外,适度的噪声注入可提升准确率,类似于随机共振效应,该效应与最优小幅值权重符号翻转相关。最后,通过二分网络随机化仅保留符号结构(而非精确权重幅值)足以维持高准确率。这些现象定义了与模型及模态无关的任务复杂性度量:全精度神经网络与二值化或随机化神经网络性能之差。我们的发现凸显了带符号二分拓扑结构在学习表征中的关键作用,并提出了与任务复杂性相适配的模型压缩与可解释性实用策略。