Today's most powerful machine learning approaches are typically designed to train stateless architectures with predefined layers and differentiable activation functions. While these approaches have led to unprecedented successes in areas such as natural language processing and image recognition, the trained models are also susceptible to making mistakes that a human would not. In this paper, we take the view that true intelligence may require the ability of a machine learning model to manage internal state, but that we have not yet discovered the most effective algorithms for training such models. We further postulate that such algorithms might not necessarily be based on gradient descent over a deep architecture, but rather, might work best with an architecture that has discrete activations and few initial topological constraints (such as multiple predefined layers). We present one attempt in our ongoing efforts to design such a training algorithm, applied to an architecture with binary activations and only a single matrix of weights, and show that it is able to form useful representations of natural language text, but is also limited in its ability to leverage large quantities of training data. We then provide ideas for improving the algorithm and for designing other training algorithms for similar architectures. Finally, we discuss potential benefits that could be gained if an effective training algorithm is found, and suggest experiments for evaluating whether these benefits exist in practice.
翻译:当今最强大的机器学习方法通常被设计用于训练具有预定义层级和可微激活函数的无状态架构。尽管这些方法在自然语言处理和图像识别等领域取得了前所未有的成功,但训练的模型也容易犯人类不会犯的错误。本文认为,真正的智能可能需要机器学习模型具备管理内部状态的能力,但我们尚未发现训练此类模型的最有效算法。我们进一步推测,这类算法未必基于深度架构的梯度下降,而可能最适合具有离散激活和较少初始拓扑约束(如多个预定义层级)的架构。我们展示了持续设计此类训练算法的一次尝试——将其应用于仅含单个权重矩阵的二进制激活架构,结果表明该算法能形成自然语言文本的有用表征,但在利用大规模训练数据方面存在局限。随后我们提出改进该算法的思路,以及为类似架构设计其他训练算法的建议。最后,我们探讨了若找到有效训练算法可能带来的潜在优势,并建议通过实验评估这些优势在实际中是否成立。