LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementation. We introduce CentaurEval, a unified, ecologically valid benchmark for measuring human-in-the-loop value in coding. CentaurEval's core innovation is its "Collaboration-Necessary" problem templates, which are intractable for standalone LLMs or humans, but solvable through effective collaboration. CentaurEval dynamically instantiates tasks from 45 templates, providing a standardized IDE for humans and a reproducible 450-task toolkit for LLMs. We benchmark 45 participants against 5 LLMs under 4 levels of human intervention. Results show that while LLMs or humans alone achieve poor pass rates (0.67% and 18.89%), human-AI collaboration significantly improves to 31.11%. Our analysis reveals an emerging co-reasoning partnership, challenging the traditional human-tool hierarchy by showing that strategic breakthroughs can originate from either humans or AI.
翻译:基于大语言模型的编码智能体正在重塑开发范式。然而,现有评估体系(无论是面向人类的传统测试还是面向大语言模型的基准测试)均未能捕捉这一变革,难以涵盖那些需要人类推理指导方案、同时依赖人工智能效率实现的问题。我们提出CentaurEval,一个统一且具有生态效度的基准测试,用于衡量编码中人在环路的价值。CentaurEval的核心创新在于其“协作必需”问题模板——这些任务对独立的大语言模型或人类均难以处理,但通过有效协作可被解决。CentaurEval从45个模板动态生成具体任务,为人类提供标准化集成开发环境,为大语言模型提供含450项任务的可复现工具集。我们在四种人类干预层级下,对45名参与者与5个大语言模型进行基准测试。结果表明,尽管单独使用大语言模型或人类通过率较低(分别为0.67%和18.89%),但人机协作显著提升至31.11%。我们的分析揭示了一种新兴的协同推理伙伴关系:战略突破既可源自人类也可源自人工智能,从而挑战了传统的人机工具层级关系。