We introduce VOCALExplore, a system designed to support users in building domain-specific models over video datasets. VOCALExplore supports interactive labeling sessions and trains models using user-supplied labels. VOCALExplore maximizes model quality by automatically deciding how to select samples based on observed skew in the collected labels. It also selects the optimal video representations to use when training models by casting feature selection as a rising bandit problem. Finally, VOCALExplore implements optimizations to achieve low latency without sacrificing model performance. We demonstrate that VOCALExplore achieves close to the best possible model quality given candidate acquisition functions and feature extractors, and it does so with low visible latency (~1 second per iteration) and no expensive preprocessing.
翻译:我们提出VOCALExplore系统,旨在支持用户在视频数据集上构建领域特定模型。该系统支持交互式标注会话,并利用用户提供的标签训练模型。VOCALExplore通过基于收集标签中观察到的偏态自动决定样本选择策略,从而最大化模型质量。此外,该系统将特征选择问题建模为增长型赌博机问题,在训练模型时自动选择最优视频表示。最后,VOCALExplore通过优化实现低延迟,同时不牺牲模型性能。我们证明,在给定候选采集函数和特征提取器的条件下,VOCALExplore能够达到近乎最优的模型质量,且具有极低的可见延迟(每次迭代约1秒),无需昂贵的预处理。