Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing frameworks are often difficult to extend and lack both interoperability and systematic verification that integrated methods faithfully reproduce their originally reported results. We introduce \emph{RecourseBench}, a unified evaluation framework built around three commitments namely, modularity, reproducibility, and interactivity. The framework decomposes the pipeline into five fully decoupled layers -- Data, Preprocessing, Model, Recourse Method, and Evaluation -- governed by abstract interfaces and a dynamic registry. To address the reproducibility gap in prior benchmarks, we introduce a four-tier classification system in which every integrated method is validated by an automated test suite against its originally reported results. We further provide an interactive web interface for flexible, configuration-driven comparison across methods, datasets, and model architectures. Our framework currently integrates 28 state-of-the-art recourse methods and, to our knowledge, constitutes the first recourse benchmark to explicitly enforce method-level reproducibility through automated, quantitative testing.
翻译:算法追责方法提供反事实解释,告知个体推翻不利模型决策所需采取的行动。尽管方法论进展迅速,但原则性比较仍难以实现——现有框架往往难以扩展,且缺乏互操作性,也无法系统验证集成方法是否忠实复现了其原始报告结果。我们提出统一评估框架\emph{RecourseBench},该框架围绕三大承诺构建:模块化、可复现性及交互性。框架将流程解耦为五个完全独立的层级——数据、预处理、模型、追责方法与评估——由抽象接口与动态注册表管控。为解决先前基准测试中的可复现性缺口,我们引入四级分类系统,通过自动化测试套件验证每项集成方法是否匹配原始报告结果。此外,我们提供交互式Web界面,支持基于配置灵活比较不同方法、数据集与模型架构。当前框架已集成28种前沿追责方法,据我们所知,这是首个通过自动化定量测试显式强制执行方法级可复现性的追责基准测试。