Machine Learning (ML) has been widely used in Natural Language Processing (NLP) applications. A fundamental assumption in ML is that training data and real-world data should follow a similar distribution. However, a deployed ML model may suffer from out-of-distribution (OOD) issues due to distribution shifts in the real-world data. Though many algorithms have been proposed to detect OOD data from text corpora, there is still a lack of interactive tool support for ML developers. In this work, we propose DeepLens, an interactive system that helps users detect and explore OOD issues in massive text corpora. Users can efficiently explore different OOD types in DeepLens with the help of a text clustering method. Users can also dig into a specific text by inspecting salient words highlighted through neuron activation analysis. In a within-subjects user study with 24 participants, participants using DeepLens were able to find nearly twice more types of OOD issues accurately with 22% more confidence compared with a variant of DeepLens that has no interaction or visualization support.
翻译:机器学习(ML)已被广泛应用于自然语言处理(NLP)应用。ML的一个基本假设是训练数据与现实世界数据应遵循相似的分布。然而,已部署的ML模型可能因现实世界数据的分布偏移而遭受分布外(OOD)问题。尽管已有许多算法被提出用于检测文本语料库中的OOD数据,但仍缺乏面向ML开发者的交互式工具支持。在本工作中,我们提出了DeepLens,一个帮助用户检测并探索海量文本语料库中OOD问题的交互式系统。借助文本聚类方法,用户可以高效探索DeepLens中的不同OOD类型。用户还可通过神经元激活分析高亮显示的关键词深入挖掘特定文本。在包含24名参与者的受试者内用户研究中,与无交互或可视化支持的DeepLens变体相比,使用DeepLens的参与者能够准确发现近两倍的OOD问题类型,且置信度提升22%。