Creating Computer Vision (CV) models remains a complex practice, despite their ubiquity. Access to data, the requirement for ML expertise, and model opacity are just a few points of complexity that limit the ability of end-users to build, inspect, and improve these models. Interactive ML perspectives have helped address some of these issues by considering a teacher in the loop where planning, teaching, and evaluating tasks take place. We present and evaluate two interactive visualizations in the context of Sprite, a system for creating CV classification and detection models for images originating from videos. We study how these visualizations help Sprite's users identify (evaluate) and select (plan) images where a model is struggling and can lead to improved performance, compared to a baseline condition where users used a query language. We found that users who had used the visualizations found more images across a wider set of potential types of model errors.
翻译:构建计算机视觉(CV)模型尽管普遍存在,但依然是一项复杂的实践。数据访问权限、机器学习专业知识需求以及模型不透明性,仅是限制终端用户构建、检查和改进这些模型能力的少数几个复杂因素。交互式机器学习视角通过引入“教师参与循环”的方式,在其中进行规划、教学和评估任务,已在一定程度上帮助解决了这些问题。我们针对Sprite系统(用于从视频图像中创建CV分类与检测模型的系统)提出并评估了两种交互式可视化方法。我们研究了与使用查询语言的基线条件相比,这些可视化方法如何帮助Sprite用户识别(评估)并选择(规划)模型难以处理且可能导致性能提升的图像。我们发现,使用可视化方法的用户能够在更广泛的潜在模型错误类型中找出更多图像。