Scientific data annotation, such as tracking animals in video or proofreading neural reconstructions, remains bottlenecked by the "last mile" problem: even with strong automation, verification and correction consume substantial human effort. Standard approaches train models to directly predict annotations, discarding the rich supervision in how experts navigate, click, verify, and correct. We introduce a framework for studying behavioral cloning on scientific annotation: 9 synthetic tasks paired with synthetic annotations that simulate realistic human strategies including exploration, mistake correction, and strategic decision-making. Our experiments reveal several findings. First, skills emerge hierarchically: models learn GUI mechanics before task-critical decisions, and commit fewer mistakes than the training data while retaining the ability to correct errors when they occur. Second, scaling models on multi-task behavioral cloning shows that larger models are more data efficient within our scale range. Third, multi-task pretraining enables efficient fine-tuning to new tasks, while training from scratch fails entirely. Fourth, linear probes reveal that models internally represent latent variables of the annotation process such as task phase and data position; interestingly, we find a shared mistake representation that generalizes across different annotation tasks. Overall, our framework establishes systematic benchmarks and identifies key bottlenecks, providing a foundation for scaling behavioral cloning to real-world scientific data annotation.
翻译:科学数据标注(如视频中的动物追踪或神经重建的校对)仍受限于“最后一英里”问题:即便具备强自动化能力,验证与修正过程仍需耗费大量人力。传统方法训练模型直接预测标注结果,却丢弃了专家在导航、点击、验证与修正过程中蕴含的丰富监督信息。我们提出面向科学标注行为克隆的研究框架:包含9个合成任务及其对应合成标注,可模拟现实人类策略(包括探索、错误修正与战略决策)。实验揭示多项发现:第一,技能呈层级式涌现——模型先习得图形用户界面操作机制,再掌握任务关键决策,其犯错率低于训练数据且具备纠错能力;第二,多任务行为克隆的规模扩展实验表明,在既定规模范围内,大模型展现出更高数据效率;第三,多任务预训练可高效微调至新任务,而从头训练则完全失败;第四,线性探测显示模型内部表征了标注过程的隐变量(如任务阶段与数据位置),尤其发现跨不同标注任务泛化的共享错误表征。总体而言,本框架建立了系统化基准并识别关键瓶颈,为行为克隆扩展至真实世界科学数据标注奠定基础。