Mobile UI understanding is important for enabling various interaction tasks such as UI automation and accessibility. Previous mobile UI modeling often depends on the view hierarchy information of a screen, which directly provides the structural data of the UI, with the hope to bypass challenging tasks of visual modeling from screen pixels. However, view hierarchies are not always available, and are often corrupted with missing object descriptions or misaligned structure information. As a result, despite the use of view hierarchies could offer short-term gains, it may ultimately hinder the applicability and performance of the model. In this paper, we propose Spotlight, a vision-only approach for mobile UI understanding. Specifically, we enhance a vision-language model that only takes the screenshot of the UI and a region of interest on the screen -- the focus -- as the input. This general architecture of Spotlight is easily scalable and capable of performing a range of UI modeling tasks. Our experiments show that our model establishes SoTA results on several representative UI tasks and outperforms previous methods that use both screenshots and view hierarchies as inputs. Furthermore, we explore multi-task learning and few-shot prompting capacities of the proposed models, demonstrating promising results in the multi-task learning direction.
翻译:移动界面理解对实现界面自动化、无障碍交互等任务至关重要。传统移动界面建模常依赖屏幕视图层级信息(直接提供界面结构数据),试图规避基于像素的视觉建模挑战。然而,视图层级并非始终可用,常出现对象描述缺失或结构信息错位等问题。尽管使用视图层级可带来短期性能提升,但最终会限制模型的适用性和表现。本文提出Spotlight方法,这是一种纯视觉的移动界面理解方案。具体而言,我们增强了一个仅以界面截图和屏幕关注区域(焦点)为输入的视觉-语言模型。该通用架构易于扩展,可执行多种界面建模任务。实验表明,本模型在多个代表性界面任务中取得最先进(SoTA)结果,性能超越同时使用截图和视图层级作为输入的现有方法。此外,我们探索了所提模型的多任务学习与少样本提示能力,在多任务学习方向上展现出显著优势。