Compared to LHC Run 1 and Run 2, future HEP experiments, e.g., at the HL-LHC, will increase the volume of generated data by an order of magnitude. In order to sustain the expected analysis throughput, ROOT's RNTuple I/O subsystem has been engineered to overcome the bottlenecks of the TTree I/O subsystem, focusing also on a compact data format, asynchronous and parallel requests, and a layered architecture that allows supporting distributed filesystem-less storage systems, e.g. HPC-oriented object stores. In a previous publication, we introduced and evaluated the RNTuple's native backend for Intel DAOS. Since its first prototype, we carried out a number of improvements both on RNTuple and its DAOS backend aiming to saturate the physical link, such as support for vector writes and an improved RNTuple-to-DAOS mapping, only to name a few. In parallel, the latest developments allow for better integration between RNTuple and ROOT's storage-agnostic, declarative interface to write HEP analyses, RDataFrame. In this work, we contribute with the following: (i) a redesign of the RNTuple DAOS backend, including a mechanism for efficient population of the object store based on existing data; and (ii) an experimental evaluation on a single-node platform, showing a significant increase in the analysis throughput for typical HEP workflows.
翻译:与LHC Run 1和Run 2相比,未来的HEP实验(例如在HL-LHC上)生成的数据量将增加一个数量级。为了维持预期的分析吞吐量,ROOT的RNTuple I/O子系统被设计用于克服TTree I/O子系统的瓶颈,同时重点关注紧凑的数据格式、异步和并行请求,以及支持无分布式文件系统存储系统(例如面向HPC的对象存储)的分层架构。在先前的研究中,我们介绍并评估了RNTuple针对Intel DAOS的原生后端。自首个原型以来,我们在RNTuple及其DAOS后端上实施了一系列改进,旨在饱和物理链路,例如支持向量写入和改进的RNTuple-to-DAOS映射,仅举几例。与此同时,最新进展使得RNTuple与ROOT的存储无关、声明式HEP分析编写接口RDataFrame之间能够更好地集成。在本文中,我们贡献了以下内容:(i)重新设计RNTuple DAOS后端,包括一种基于现有数据高效填充对象存储的机制;以及(ii)在单节点平台上进行的实验评估,显示典型HEP工作流的分析吞吐量显著提升。