We present AutoDOViz, an interactive user interface for automated decision optimization (AutoDO) using reinforcement learning (RL). Decision optimization (DO) has classically being practiced by dedicated DO researchers where experts need to spend long periods of time fine tuning a solution through trial-and-error. AutoML pipeline search has sought to make it easier for a data scientist to find the best machine learning pipeline by leveraging automation to search and tune the solution. More recently, these advances have been applied to the domain of AutoDO, with a similar goal to find the best reinforcement learning pipeline through algorithm selection and parameter tuning. However, Decision Optimization requires significantly more complex problem specification when compared to an ML problem. AutoDOViz seeks to lower the barrier of entry for data scientists in problem specification for reinforcement learning problems, leverage the benefits of AutoDO algorithms for RL pipeline search and finally, create visualizations and policy insights in order to facilitate the typical interactive nature when communicating problem formulation and solution proposals between DO experts and domain experts. In this paper, we report our findings from semi-structured expert interviews with DO practitioners as well as business consultants, leading to design requirements for human-centered automation for DO with RL. We evaluate a system implementation with data scientists and find that they are significantly more open to engage in DO after using our proposed solution. AutoDOViz further increases trust in RL agent models and makes the automated training and evaluation process more comprehensible. As shown for other automation in ML tasks, we also conclude automation of RL for DO can benefit from user and vice-versa when the interface promotes human-in-the-loop.
翻译:我们提出AutoDOViz,一种利用强化学习(RL)实现自动决策优化(AutoDO)的交互式用户界面。决策优化(DO)传统上由专业DO研究人员实践,专家需要花费大量时间通过试错法微调解决方案。AutoML流水线搜索旨在通过利用自动化搜索和调优解决方案,使数据科学家更容易找到最佳机器学习流水线。近期,这些进展被应用于AutoDO领域,其类似目标是通过算法选择和参数调优找到最佳强化学习流水线。然而,与机器学习问题相比,决策优化需要更复杂的问题规范。AutoDOViz旨在降低数据科学家在强化学习问题规范方面的入门门槛,利用AutoDO算法优势进行RL流水线搜索,最终创建可视化和策略洞察,以促进DO专家与领域专家在沟通问题表述和解决方案提案时典型的交互特性。本文报告了我们对DO从业者及业务顾问进行的半结构化专家访谈结果,据此提出基于强化学习的人本自动决策优化设计需求。我们通过数据科学家对系统实现的评估发现,在采用我们的方案后,他们参与决策优化的意愿显著增强。AutoDOViz进一步提升了用户对RL智能体模型的信任度,使自动化训练和评估过程更易于理解。与机器学习任务中其他自动化案例类似,我们同样得出结论:当界面支持人在回环机制时,RL驱动决策优化的自动化可实现用户与系统的互利共赢。