High-performance computing (HPC) systems are increasingly exploring dynamic resource management and malleable MPI applications to better adapt to heterogeneous architectures, fluctuating workloads, and energy constraints. However, the correctness of the libraries that support these techniques is often evaluated through ad hoc experiments that can be difficult to reproduce and maintain. This article introduces methodology for testing dynamic resource management frameworks that combines a taxonomy of tests for MPI malleable libraries with an HPC-oriented continuous integration (CI) ecosystem. The taxonomy structures functional and non-functional tests at both component-integration and system levels. The CI ecosystem instantiates this taxonomy in a containerized virtual cluster enabling automated validation. The approach is instantiated and evaluated using the Dynamic Management of Resources (DMR) framework as a representative case study. Results show that the proposed methodology improves early fault detection, simplifies maintenance under evolving dependencies, and transfers to other malleability solutions that expose analogous primitives for initialization, readiness checking, and reconfiguration.
翻译:高性能计算(HPC)系统正越来越多地探索动态资源管理和可塑性MPI应用,以更好地适应异构架构、波动的负载和能源约束。然而,支持这些技术的库的正确性通常通过难以复现和维护的特设实验进行评估。本文介绍了一种用于测试动态资源管理框架的方法论,该方法将MPI可塑性库的测试分类与面向HPC的持续集成(CI)生态系统相结合。该分类法在组件集成和系统层面上结构化功能性与非功能性测试。CI生态系统在容器化虚拟集群中实例化此分类法,实现自动化验证。该方案以动态资源管理(DMR)框架为典型案例进行实例化和评估。结果表明,所提出的方法论能改进早期故障检测,简化依赖关系演化下的维护工作,并能迁移至其他暴露类似初始化、就绪性检查和重配置原语的可塑性解决方案。