The burgeoning growth of public domain data and the increasing complexity of deep learning model architectures have underscored the need for more efficient data representation and analysis techniques. This paper is motivated by the work of (Helal, 2023) and aims to present a comprehensive overview of tensorization. This transformative approach bridges the gap between the inherently multidimensional nature of data and the simplified 2-dimensional matrices commonly used in linear algebra-based machine learning algorithms. This paper explores the steps involved in tensorization, multidimensional data sources, various multiway analysis methods employed, and the benefits of these approaches. A small example of Blind Source Separation (BSS) is presented comparing 2-dimensional algorithms and a multiway algorithm in Python. Results indicate that multiway analysis is more expressive. Contrary to the intuition of the dimensionality curse, utilising multidimensional datasets in their native form and applying multiway analysis methods grounded in multilinear algebra reveal a profound capacity to capture intricate interrelationships among various dimensions while, surprisingly, reducing the number of model parameters and accelerating processing. A survey of the multi-away analysis methods and integration with various Deep Neural Networks models is presented using case studies in different application domains.
翻译:公共领域数据的迅猛增长以及深度学习模型架构日益复杂,凸显了对更高效数据表示与分析技术的迫切需求。本文受 Helal(2023)工作的启发,旨在系统综述张量化方法。这一变革性技术弥合了数据固有的多维本质与线性代数机器学习算法中常用的简化二维矩阵之间的鸿沟。本文探讨了张量化过程中的关键步骤、多维数据来源、所采用的各种多路分析方法及其优势。通过一个盲源分离(BSS)的小型实例,在 Python 环境中对比了二维算法与多路算法的性能差异。结果表明多路分析方法具有更强的表达能力。与维度灾难的直觉相反,在保持多维数据集原生形态的基础上,采用基于多重线性代数的多路分析方法,不仅能捕捉不同维度间复杂的内在关联,更令人意外地降低了模型参数数量并加快了处理速度。本文通过多路分析方法与各类深度神经网络模型融合的应用案例研究,系统综述了不同领域的典型应用场景。