The growing adoption of RISC-V in high-performance and scientific computing has increased the need for performance-portable code targeting the RISC-V Vector (RVV) extension. However, current compiler infrastructures provide limited end-to-end support for generating optimized RVV code from high-level representations to low-level implementations. In particular, existing MLIR distributions lack practical lowering paths that map high-level abstractions to RVV intrinsics, limiting their applicability for production-ready RISC-V kernels. This paper presents a compilation approach that combines MLIR with xDSL to bridge the missing lowering stages required for RVV code generation. Using custom intermediate representations and transformation passes implemented in xDSL, we systematically translate high-level operations into specialized, hardware-aware C code invoking RVV intrinsics. The resulting kernels are emitted as portable C functions that can be directly integrated into existing applications, enabling incremental adoption without modifying surrounding software stacks. We demonstrate the approach on the General Matrix Multiplication (GEMM) kernel and evaluate the generated micro-kernels on two real RISC-V platforms, the K230 and the BananaPi F3, comparing against OpenBLAS for both square-matrix benchmarks and transformer-based workloads derived from the BERT-Large model. When integrated into a matrix multiplication kernel, the proposed approach consistently outperforms OpenBLAS, reaching up to 12.2 GFLOPS compared to the baseline's 5.1 GFLOPS and providing performance improvements between 10--35\% across the evaluated workloads. These results demonstrate that combining MLIR with xDSL provides a practical pathway to portable, optimized code generation for RISC-V platforms.


翻译:随着RISC-V在高性能与科学计算领域的广泛采用,面向RISC-V向量(RVV)扩展的可移植性能代码需求日益增长。然而,当前编译器基础设施对从高层表示到低层实现的优化RVV代码生成,仅提供有限的端到端支持。具体而言,现有MLIR发行版缺乏将高层抽象映射至RVV内建函数(intrinsics)的实用降级路径,这限制了其在生产级RISC-V内核中的应用。本文提出一种结合MLIR与xDSL的编译方法,以弥合RVV代码生成所需的缺失降级阶段。通过使用xDSL实现的自定义中间表示与变换通道,我们系统性地将高层操作转换为调用RVV内建函数的专用硬件感知C代码。生成的内核以可移植C函数形式输出,可直接集成至现有应用程序中,从而在不修改周边软件栈的情况下实现渐进式采用。我们以通用矩阵乘法(GEMM)内核演示该方法,并在两个真实的RISC-V平台(K230与BananaPi F3)上对生成的微内核进行评估,在方阵基准测试及基于BERT-Large模型的Transformer负载中均与OpenBLAS进行对比。集成至矩阵乘法内核后,所提方法始终优于OpenBLAS,在基线仅达5.1 GFLOPS的场景下实现高达12.2 GFLOPS的性能,并在全部评估负载中带来10%至35%的性能提升。这些结果表明,MLIR与xDSL的结合为面向RISC-V平台的可移植优化代码生成提供了一条实用路径。

0
下载
关闭预览

相关内容

代码(Code)是专知网的一个重要知识资料文档板块,旨在整理收录论文源代码、复现代码,经典工程代码等,便于用户查阅下载使用。
通过强化学习增强代码生成中的代码大语言模型:综述
专知会员服务
30+阅读 · 2025年1月1日
《用于代码弱点识别的 LLVM 中间表示》CMU
专知会员服务
15+阅读 · 2022年12月12日
赛尔笔记 | 条件变分自编码器(CVAE)
AINLP
28+阅读 · 2019年11月8日
基于R语言进行Box-Cox变换
R语言中文社区
45+阅读 · 2018年11月19日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 4月13日
VIP会员
最新内容
俄乌战争中关于中程打击无人机部署的经验启示
专知会员服务
0+阅读 · 17分钟前
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
相关VIP内容
通过强化学习增强代码生成中的代码大语言模型:综述
专知会员服务
30+阅读 · 2025年1月1日
《用于代码弱点识别的 LLVM 中间表示》CMU
专知会员服务
15+阅读 · 2022年12月12日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员