LLM-powered computer-use agents (CUAs) are shifting users from direct manipulation to supervisory coordination. Existing oversight mechanisms, however, have largely been studied as isolated interface features, making broader oversight strategies difficult to compare. We conceptualize CUA oversight as a structural coordination problem defined by delegation structure and engagement level, and use this lens to compare four oversight strategies in a mixed-methods study with 48 participants in a live web environment. Our results show that oversight strategy more reliably shaped users' exposure to problematic actions than their ability to correct them once visible. Plan-based strategies were associated with lower rates of agent problematic-action occurrence, but not equally strong gains in runtime intervention success once such actions became visible. On subjective measures, no single strategy was uniformly best, and the clearest context-sensitive differences appeared in trust. Qualitative findings further suggest that intervention depended not only on what controls users retained, but on whether risky moments became legible as requiring judgment during execution. These findings suggest that effective CUA oversight is not achieved by maximizing human involvement alone. Instead, it depends on how supervision is structured to surface decision-critical moments and support their recognition in time for meaningful intervention.
翻译:大型语言模型驱动的计算机使用智能体(CUA)正将用户从直接操作转向监督式协调。然而,现有监督机制大多作为孤立的界面特征进行研究,使得更广泛的监督策略难以相互比较。我们将CUA监督概念化为一个由授权结构和参与程度定义的结构性协调问题,并借助这一视角,在一项混合方法研究中,于真实网络环境中对48名参与者比较了四种监督策略。结果表明,监督策略在塑造用户接触问题行为方面的作用,比用户在问题行为显现后纠正它们的能力更为稳定。基于计划的策略与智能体问题行为发生率较低相关,但在行为显现后,其在运行时干预成功率方面的提升并不同样显著。在主观测量方面,没有任何一种策略普遍最优,而最明显的上下文敏感差异体现在信任度上。定性研究结果进一步表明,干预不仅取决于用户保留了哪些控制权,还取决于高风险时刻在执行过程中是否能被识别为需要判断的节点。这些发现表明,有效的CUA监督并非仅通过最大化人工参与即可实现,而是取决于如何设计监督结构,以凸显决策关键时刻,并支持用户及时识别这些时刻,从而进行有意义的干预。