分布式异构张量-向量收缩算法在高性能计算中的应用 (Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing) - 专知论文

会员服务 ·

0

Performer · 收缩 · TVC · Tensor · 流 ·

2025 年 1 月 6 日

Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

翻译：分布式异构张量-向量收缩算法在高性能计算中的应用

Pedro J. Martinez-Ferrer,Albert-Jan Yzelman,Vicenç Beltran

from arxiv, 15 pages, 7 figures, preprint (accepted for publication at Journal of Future Generation Computer Systems)

The tensor-vector contraction (TVC) is the most memory-bound operation of its class and a core component of the higher order power method (HOPM). This paper brings distributed-memory parallelization to a native TVC algorithm for dense tensors that overall remains oblivious to contraction mode, tensor splitting and tensor order. Similarly, we propose a novel distributed HOPM, namely dHOPM3, that can save up to one order of magnitude of streamed memory and is about twice as costly in terms of data movement as a distributed TVC operation (dTVC) when using task-based parallelization. The numerical experiments carried out in this work on three different architectures featuring multi-core and accelerated systems confirm that the performance of dTVC and dHOPM3 remains relatively close to the peak system memory bandwidth (50%-80%, depending on the architecture) and on par with STREAM reference values. On strong scalability scenarios, our native multi-core implementations of these two algorithms can achieve similar and sometimes even greater performance figures than those based upon state-of-the-art CUDA batched kernels. Finally, we demonstrate that both computation and communication can benefit from mixed precision arithmetic also in cases where the hardware does not support low precision data types natively.

翻译：张量-向量收缩（TVC）是该类运算中内存受限最严重的操作，也是高阶幂法（HOPM）的核心组成部分。本文为稠密张量的原生TVC算法引入了分布式内存并行化，该算法整体上对收缩模式、张量分割和张量阶数保持透明。相应地，我们提出了一种新颖的分布式HOPM算法——dHOPM3，该算法可节省高达一个数量级的流式内存访问量，且在使用基于任务的并行化时，其数据移动开销约为分布式TVC操作（dTVC）的两倍。本研究在多核及加速器系统构成的三种不同架构上进行的数值实验表明，dTVC与dHOPM3的性能始终接近系统内存带宽峰值（50%-80%，具体取决于架构），并与STREAM基准测试值相当。在强可扩展性场景中，我们针对这两种算法的原生多核实现，其性能指标与基于前沿CUDA批处理内核的实现相当，有时甚至更优。最后，我们证明了即使在硬件本身不支持低精度数据类型的情况下，混合精度运算仍能同时提升计算与通信效率。

0

相关内容

Performer

【普林斯顿博士论文】图机器学习，137页pdf

【普林斯顿博士论文】图机器学习，137页pdf

专知会员服务

43+阅读 · 2024年5月1日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日

【ACL2020】多模态信息抽取，365页ppt

【ACL2020】多模态信息抽取，365页ppt

专知会员服务

151+阅读 · 2020年7月6日

自动结构变分推理，Automatic structured variational inference

自动结构变分推理，Automatic structured variational inference

专知会员服务

41+阅读 · 2020年2月10日

【AI应用】Facebook-利用神经网络求解高等数学方程, Using neural networks to solve advanced mathematics equations

【AI应用】Facebook-利用神经网络求解高等数学方程, Using neural networks to solve advanced mathematics equations

专知会员服务

34+阅读 · 2020年1月15日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

学习自然语言处理路线图

学习自然语言处理路线图

专知会员服务

140+阅读 · 2019年9月24日

Transformer模型-深度学习自然语言处理，17页ppt

Transformer模型-深度学习自然语言处理，17页ppt

专知

14+阅读 · 2020年8月30日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

利用动态深度学习预测金融时间序列基于Python

利用动态深度学习预测金融时间序列基于Python

量化投资与机器学习

18+阅读 · 2018年10月30日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

开放知识图谱

28+阅读 · 2018年6月11日

CNN 反向传播算法推导

CNN 反向传播算法推导

统计学习与视觉计算组

30+阅读 · 2017年12月29日

基于位置注意力机制模型和带标签数据来提升槽填充（EMNLP outstanding paper）

基于位置注意力机制模型和带标签数据来提升槽填充（EMNLP outstanding paper）

科技创新与创业

17+阅读 · 2017年11月17日

基于LDA的主题模型实践（三）

基于LDA的主题模型实践（三）

机器学习深度学习实战原创交流

23+阅读 · 2015年10月12日

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

Musielak-Orlicz-Sobolev 空间中的迹嵌入及其应用

国家自然科学基金

2+阅读 · 2015年12月31日

直接优化半周长线长的VLSI两阶段迭代布局算法研究

国家自然科学基金

0+阅读 · 2015年12月31日

三类多尺度问题的多尺度算法

国家自然科学基金

1+阅读 · 2015年12月31日

Schr？dinger-Poisson方程守恒DDG方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

随机二阶锥互补问题理论与算法研究及其应用

国家自然科学基金

0+阅读 · 2015年12月31日

动态Gr？bner 基与GVW算法

国家自然科学基金

0+阅读 · 2014年12月31日

Biot模型基于有限元离散的多重网格算法研究

国家自然科学基金

1+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

关于二阶锥互补约束数学规划问题的约束规范和算法研究

国家自然科学基金

0+阅读 · 2014年12月31日

Infinite-dimensional next-generation reservoir computing

Arxiv

0+阅读 · 2025年2月21日

A robust estimation and variable selection approach for sparse partially linear additive models

Arxiv

0+阅读 · 2025年2月18日

A numerical method for solving the generalized tangent vector of hyperbolic systems

Arxiv

0+阅读 · 2025年2月17日

Optimal design of experiments with quantitative-sequence factors

Arxiv

0+阅读 · 2025年2月5日

Sensitivity analysis for multivariable missing data using multiple imputation: a tutorial

Arxiv

0+阅读 · 2025年2月5日

A hybrid numerical method for elastic wave propagation in discontinuous media with complex geometry

Arxiv

0+阅读 · 2025年2月3日

Using gradient of Lagrangian function to compute efficient channels for the ideal observer

Arxiv

0+阅读 · 2025年1月31日

Modernizing full posterior inference for surrogate modeling of categorical-output simulation experiments

Arxiv

0+阅读 · 2025年1月24日

Parallel remote state preparation for fully device-independent verifiable blind quantum computation

Arxiv

0+阅读 · 2025年1月22日

Module-conditioned distribution of quantum circuits

Arxiv

0+阅读 · 2025年1月21日

VIP会员

文章信息

相关主题

相关VIP内容

【普林斯顿博士论文】图机器学习，137页pdf

【普林斯顿博士论文】图机器学习，137页pdf

专知会员服务

43+阅读 · 2024年5月1日

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日

【ACL2020】多模态信息抽取，365页ppt

【ACL2020】多模态信息抽取，365页ppt

专知会员服务

151+阅读 · 2020年7月6日

自动结构变分推理，Automatic structured variational inference

自动结构变分推理，Automatic structured variational inference

专知会员服务

41+阅读 · 2020年2月10日

【AI应用】Facebook-利用神经网络求解高等数学方程, Using neural networks to solve advanced mathematics equations

【AI应用】Facebook-利用神经网络求解高等数学方程, Using neural networks to solve advanced mathematics equations

专知会员服务

34+阅读 · 2020年1月15日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

学习自然语言处理路线图

学习自然语言处理路线图

专知会员服务

140+阅读 · 2019年9月24日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】基于自适应表征的高效视觉建模

《多域作战中融合网络、电子战与动能机动》

AI智能体时代大模型安全风险与攻防新挑战

迈向个性化大语言模型驱动的智能体：基础、评估与未来方向

相关资讯

Transformer模型-深度学习自然语言处理，17页ppt

Transformer模型-深度学习自然语言处理，17页ppt

专知

14+阅读 · 2020年8月30日

基于深度元学习的因果推断新方法

基于深度元学习的因果推断新方法

图与推荐

12+阅读 · 2020年7月21日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

利用动态深度学习预测金融时间序列基于Python

利用动态深度学习预测金融时间序列基于Python

量化投资与机器学习

18+阅读 · 2018年10月30日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

论文浅尝 | 嵌入常识知识的注意力 LSTM 模型用于特定目标的基于侧面的情感分析

开放知识图谱

28+阅读 · 2018年6月11日

CNN 反向传播算法推导

CNN 反向传播算法推导

统计学习与视觉计算组

30+阅读 · 2017年12月29日

基于位置注意力机制模型和带标签数据来提升槽填充（EMNLP outstanding paper）

基于位置注意力机制模型和带标签数据来提升槽填充（EMNLP outstanding paper）

科技创新与创业

17+阅读 · 2017年11月17日

基于LDA的主题模型实践（三）

基于LDA的主题模型实践（三）

机器学习深度学习实战原创交流

23+阅读 · 2015年10月12日

相关论文

Infinite-dimensional next-generation reservoir computing

Arxiv

0+阅读 · 2025年2月21日

A robust estimation and variable selection approach for sparse partially linear additive models

Arxiv

0+阅读 · 2025年2月18日

A numerical method for solving the generalized tangent vector of hyperbolic systems

Arxiv

0+阅读 · 2025年2月17日

Optimal design of experiments with quantitative-sequence factors

Arxiv

0+阅读 · 2025年2月5日

Sensitivity analysis for multivariable missing data using multiple imputation: a tutorial

Arxiv

0+阅读 · 2025年2月5日

A hybrid numerical method for elastic wave propagation in discontinuous media with complex geometry

Arxiv

0+阅读 · 2025年2月3日

Using gradient of Lagrangian function to compute efficient channels for the ideal observer

Arxiv

0+阅读 · 2025年1月31日

Modernizing full posterior inference for surrogate modeling of categorical-output simulation experiments

Arxiv

0+阅读 · 2025年1月24日

Parallel remote state preparation for fully device-independent verifiable blind quantum computation

Arxiv

0+阅读 · 2025年1月22日

Module-conditioned distribution of quantum circuits

Arxiv

0+阅读 · 2025年1月21日

相关基金

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

Musielak-Orlicz-Sobolev 空间中的迹嵌入及其应用

国家自然科学基金

2+阅读 · 2015年12月31日

直接优化半周长线长的VLSI两阶段迭代布局算法研究

国家自然科学基金

0+阅读 · 2015年12月31日

三类多尺度问题的多尺度算法

国家自然科学基金

1+阅读 · 2015年12月31日

Schr？dinger-Poisson方程守恒DDG方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

随机二阶锥互补问题理论与算法研究及其应用

国家自然科学基金

0+阅读 · 2015年12月31日

动态Gr？bner 基与GVW算法

国家自然科学基金

0+阅读 · 2014年12月31日

Biot模型基于有限元离散的多重网格算法研究

国家自然科学基金

1+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

关于二阶锥互补约束数学规划问题的约束规范和算法研究

国家自然科学基金

0+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员