FL systems are inherently subject to client heterogeneity arising from differences in hardware capabilities. We propose a realistic evaluation framework for hardware-aware federated learning methods based on lightweight emulation of client hardware. Existing evaluation approaches address this either through small-scale real-device deployments or through trace-driven and probabilistic simulations. The former are difficult to scale and reproduce, while the latter suffer from three compounding sources of uncertainty: the choice of distribution family, the choice of execution-time estimates, and the inability to capture workload-dependent behavior. Our framework reproduces heterogeneous client behavior at execution time, capturing compute capacity, memory-constrained feasibility, and their interaction within a single host system. Unlike fixed-trace or throughput-scaling approaches, the proposed method is workload-aware: the same client population can exhibit different runtime and failure behavior depending on the task under study. We evaluate the fidelity of the approach across multiple workloads and show it preserves relative device performance while accurately reflecting hardware-dependent execution constraints. Compared against direct hardware measurements and benchmark references, the emulation accurately reproduces feasibility and training performance, preserving both relative device ordering and absolute compute times. We demonstrate the importance of workload-aware compute times by evaluating different heterogeneity-management methods across multiple workloads. By coupling emulation with real-world-based device sampling, our framework enables realistic, scalable, and reproducible evaluation of federated learning systems under heterogeneous learning conditions, while providing a practical way to generate workload-specific runtime behavior for large and diverse client populations.
翻译:暂无翻译