Recent human-computer interaction (HCI) research has revealed a widespread misalignment between how developers design workplace artificial intelligence (AI) systems, and what workers actually need from them. Yet, little research has examined the effects of this gap, or how it may cause harm. We analyzed 1,524 reports of incidents in which AI systems were used to perform 171 occupational tasks across 12 industry sectors. Using an Large Language Model (LLM)-as-an-expert approach, we extracted the main traits of the AI systems involved in those incidents using an established framework of twelve traits. We then compared them with the traits that 202 workers highly familiar with those tasks would have preferred. We found that as many as 83\% of workplace incidents stem from worker-AI misalignments. In most cases, workers wanted systems that are precise, insightful, or personal, but instead received systems that are basic, simple, or general. Over the years, fast AI caused a considerable number of incidents, yet these declined, and imaginative AI, with the mass introduction of generative AI, started to cause incidents. We also compared the traits causing the incidents with the traits that 197 developers building AI systems for those tasks would have preferred. If the traits causing the incidents were the same as those designed by developers, then developers may be responsible for those incidents. We found that 74\% of task misalignments could be attributed to developers who tended to overfocus on efficiency and speed, especially for systems performing tasks in people-facing occupations such as those in the human resources sector. Our results call for design interventions that better align AI development with workers' needs, as without such corrections, workplace AI incidents are likely to persist, causing the invisible erosion of worker agency and organizational productivity.
翻译:近期人机交互(HCI)研究发现,开发者设计工作场所人工智能(AI)系统的方式与员工实际需求之间存在普遍脱节。然而,这种差异的影响及其可能造成的危害鲜有研究涉及。我们分析了1,524份事故报告,这些事故涉及AI系统在12个行业领域内执行171种职业任务。采用基于大语言模型(LLM)的专家评估方法,我们利用既定十二特征框架提取了涉事AI系统的主要特征,并将这些特征与202名高度熟悉相关任务的员工所偏好的特征进行对比。研究发现高达83%的工作场所事故源于员工与AI系统的特征错配。在多数案例中,员工期望获得精准、深刻或个性化系统,但实际部署的却是基础、简单或通用型系统。随时间推移,快速AI系统曾导致大量事故(现已减少),而伴随生成式AI的普及,想象型AI开始引发事故。我们进一步将事故诱因特征与197名相关AI系统开发者偏好的特征进行对比——若两者一致,则开发者可能需对事故负责。结果显示74%的任务错配可归因于开发者过度关注效率与速度,尤其在面向人群的职业领域(如人力资源部门系统)表现显著。本研究呼吁通过设计干预措施使AI开发与员工需求更紧密结合,否则工作场所AI事故将持续发生,导致员工自主权与组织生产力的隐形侵蚀。