AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that miss how developers actually experience misalignment. We present an observational study of 20,574 coding-agent sessions from 1,639 repositories across IDE and CLI workflows. We operationalize misalignment as a breakdown made visible through developer pushback, and annotate each episode along four axes: form, cause, cost, and resolution. We identify seven recurring forms, spanning how agents read projects, interpret developer intent, follow rules, bound their actions, implement and execute code, and report progress. 90.50\% of episodes impose effort and trust costs rather than irreversible system damage, yet 91.49\% of visible resolutions still require explicit user correction. Misalignment patterns also differ across IDE and CLI settings, persist across adjacent sessions, and shift over time: while overall rates decline, constraint violations and inaccurate self-reporting grow in share. Our findings inform the design of training, evaluation, and interfaces for keeping coding agents aligned with real developer workflows.
翻译:人工智能编码代理日益直接在软件环境中运作,然而对其失败方式的现有分析依赖于基准测试轨迹,未能捕捉开发者实际体验到的代理行为不协调。我们针对跨IDE和CLI工作流的1,639个代码仓库中的20,574个编码代理会话进行了一项观察性研究。我们将“不协调”定义为通过开发者反馈可见的代理行为崩溃,并沿四个维度(形式、原因、成本与解决方式)对每个事件进行标注。我们识别出七种常见的不协调形式,涵盖代理在读取项目、理解开发者意图、遵守规则、界定自身行为边界、实现与执行代码以及汇报进展等方面的表现。90.50%的事件主要造成信任与努力成本而非不可逆的系统性损害,且91.49%的可见解决方式仍需开发者主动纠正。此外,不协调模式在IDE与CLI环境中存在差异,在相邻会话间持续存在,并随时间动态变化:尽管整体发生率下降,但约束违反与不准确自我报告的比例却在上升。本研究成果为设计训练、评估与交互界面提供指导,以确保编码代理与真实开发工作流的协调一致。