Virtual assistants have been widely used by mobile phone users in recent years. Although their capabilities of processing user intents have been developed rapidly, virtual assistants in most platforms are only capable of handling pre-defined high-level tasks supported by extra manual efforts of developers. However, instance-level user intents containing more detailed objectives with complex practical situations, are yet rarely studied so far. In this paper, we explore virtual assistants capable of processing instance-level user intents based on pixels of application screens, without the requirements of extra extensions on the application side. We propose a novel cross-modal deep learning pipeline, which understands the input vocal or textual instance-level user intents, predicts the targeting operational area, and detects the absolute button area on screens without any metadata of applications. We conducted a user study with 10 participants to collect a testing dataset with instance-level user intents. The testing dataset is then utilized to evaluate the performance of our model, which demonstrates that our model is promising with the achievement of 64.43% accuracy on our testing dataset.
翻译:近年来,虚拟助手已被智能手机用户广泛采用。尽管其处理用户意图的能力发展迅速,但大多数平台的虚拟助手仅能处理由开发者额外人工支持预定义的高层任务。然而,包含更详细目标与复杂实际场景的实例级用户意图至今鲜有研究。本文探索了一种无需对应用端进行额外扩展,基于应用屏幕像素处理实例级用户意图的虚拟助手。我们提出了一种新颖的跨模态深度学习流水线,该流水线能理解输入语音或文本形式的实例级用户意图,预测目标操作区域,并在无任何应用元数据的情况下检测屏幕上的绝对按钮区域。我们开展了包含10名参与者的用户研究,收集了包含实例级用户意图的测试数据集。该测试数据集随后用于评估模型性能,结果表明,我们的模型在测试数据集上达到了64.43%的准确率,展现了良好的应用前景。