Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in the environment (e.g., README files, code comments, stack traces) and determine their relevance to the task. This creates a fundamental challenge: relevant cues must be followed to complete a task, whereas irrelevant or misleading ones must be ignored. Existing benchmarks do not capture this ability. An agent may appear capable by blindly following all instructions, or appear robust by ignoring them altogether. We introduce TAB (Task Alignment Benchmark), a suite of 89 terminal tasks derived from Terminal-Bench 2.1. Each task is intentionally underspecified, with missing information provided as a necessary cue embedded in a natural environmental artifact, alongside a plausible but irrelevant distractor. Solving these tasks requires selectively using the cue while ignoring the distractor. Applying TAB to ten frontier agents reveals a systematic gap between task capability and task alignment. The strongest Terminal-Bench agent achieves high task completion but low task alignment on TAB. Evaluating six prompt-injection defenses further shows that suppressing distractor execution also suppresses the cues required for task completion. These results demonstrate that task-aligned agents require selective use of environmental instructions rather than blanket acceptance or rejection.
翻译:终端智能体正日益能够仅凭单一用户指令自主执行复杂的长周期任务。为此,它们必须解读环境中遇到的信息(例如 README 文件、代码注释、堆栈跟踪),并判断这些信息与当前任务的相关性。这带来了一个根本性挑战:必须遵循相关线索以完成任务,而无关或误导性信息则必须被忽略。现有基准测试未能捕捉这一能力。一个智能体可能因盲目遵循所有指令而显得能力强,或因完全忽略所有指令而显得鲁棒。我们提出了 TAB(任务对齐基准),这是一套源自 Terminal-Bench 2.1 的 89 个终端任务。每个任务都故意描述不充分,缺失的信息以必要线索的形式嵌入在自然环境的产物中,同时附带一个貌似合理但无关的干扰项。解决这些任务需要选择性使用线索而忽略干扰项。将 TAB 应用于十个前沿智能体,揭示了任务能力与任务对齐之间的系统性差距。在 Terminal-Bench 上最强的智能体在 TAB 上表现出高任务完成度但低任务对齐度。进一步评估六种提示注入防御方法表明,抑制干扰项的执行也会抑制完成任务所需的线索。这些结果证明,任务对齐的智能体需要选择性使用环境指令,而非全盘接受或全盘拒绝。