Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files. However, many real-world workflows require practitioners to work directly with audio and video files. Working with such multimedia files calls for terminal agents not only to understand multimedia content, but also to convert auditory and visual evidence across related files into appropriate actions. To evaluate terminal agents on multimedia-file tasks, we introduce MultiMedia-TerminalBench (MMTB), a benchmark of 105 tasks across 5 meta-categories where terminal agents directly operate with audio and video files. Alongside MMTB, we propose Terminus-MM, a multimedia harness that extends Terminus-KIRA with audio and video perception for terminal agents. Together, MMTB and Terminus-MM support a controlled study of multimedia terminal agents, revealing how different forms of multimedia access shape task outcomes and determine which evidence agents rely on to construct executable terminal workflows. MMTB media and metadata are released at https://huggingface.co/datasets/mm-tbench/mmtb-media
翻译:终端通过提供多样化的工具来自动化复杂工作流,为AI代理提供了强大的交互界面,然而现有的终端代理基准测试主要聚焦于文本、代码和结构化文件任务。实际上,许多真实工作流要求操作人员直接处理音频和视频文件。处理此类多媒体文件不仅需要终端代理理解多媒体内容,还需将跨相关文件的听觉与视觉证据转化为恰当的执行动作。为评估终端代理在多媒体文件任务中的表现,我们提出MultiMedia-TerminalBench(MMTB)基准,该基准包含5个元类别的105项任务,使终端代理直接操作音频与视频文件。配合MMTB,我们提出Terminus-MM多媒体适配框架,该框架在Terminus-KIRA基础上扩展了针对终端代理的音频与视频感知能力。MMTB与Terminus-MM共同支持对多媒体终端代理的受控研究,揭示了不同多媒体访问形式如何影响任务结果,并决定了代理依赖哪些证据构建可执行的终端工作流。MMTB媒体及元数据发布于https://huggingface.co/datasets/mm-tbench/mmtb-media