The recently proposed training-free NAS methods abandon the training phase and design various zero-cost proxies as scores to identify excellent architectures, arousing extreme computational efficiency for neural architecture search. In this paper, we raise an interesting problem: can we properly measure the operation importance in DARTS through a training-free way, with avoiding the parameter-intensive bias? We investigate this question through the lens of edge connectivity, and provide an affirmative answer by defining a connectivity concept, ZERo-cost Operation Sensitivity (ZEROS), to score the importance of candidate operations in DARTS at initialization. By devising an iterative and data-agnostic manner in utilizing ZEROS for NAS, our novel trial leads to a framework called training free differentiable architecture search (FreeDARTS). Based on the theory of Neural Tangent Kernel (NTK), we show the proposed connectivity score provably negatively correlated with the generalization bound of DARTS supernet after convergence under gradient descent training. In addition, we theoretically explain how ZEROS implicitly avoids parameter-intensive bias in selecting architectures, and empirically show the searched architectures by FreeDARTS are of comparable size. Extensive experiments have been conducted on a series of search spaces, and results have demonstrated that FreeDARTS is a reliable and efficient baseline for neural architecture search.
翻译:近期提出的免训练神经架构搜索方法摒弃了训练阶段,通过设计各类零成本代理指标作为评分来识别优秀架构,从而极大提升了神经架构搜索的计算效率。本文提出一个有趣的问题:能否通过免训练方式合理衡量DARTS中操作的重要性,同时避免参数密集型偏差?我们从边连接性的视角探究该问题,通过定义连接性概念——零成本操作灵敏度(ZEROS),在初始化阶段对DARTS候选操作的重要性进行评分。通过设计一种迭代且与数据无关的方式将ZEROS应用于神经架构搜索,我们提出名为"免训练可微架构搜索(FreeDARTS)"的新框架。基于神经正切核理论,我们证明了所提出的连接性评分与梯度下降训练收敛后DARTS超网络的泛化界负相关。此外,我们从理论上解释了ZEROS如何隐式避免架构选择中的参数密集型偏差,并实验表明FreeDARTS搜索得到的架构具有可比规模。在多个搜索空间上进行了广泛实验,结果表明FreeDARTS是神经架构搜索中可靠且高效的基准方法。