Despite recent progress in language models and agents for scientific data-driven discovery, further advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we introduce D3-Gym, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Discovery. D3-Gym comprises (1) 565 tasks sourced from 239 real scientific repositories across four disciplines where (2) each task is equipped with a natural language instruction, an executable environment with pre-installed dependencies, input dataset and artifact previews, a reference code solution, and an automatically synthesized evaluation script. Rigorous evaluation of the quality of the verification signal in D3-Gym confirms that our evaluation scripts achieve 87.5% agreement with human-annotated gold standards and strong alignment in domain-specific evaluation logic, showing their scientific soundness. Further, training on trajectories sampled from D3-Gym yields consistent and substantial gains across Qwen3 models of varying sizes on ScienceAgentBench, boosting Qwen3-32B by 7.8 absolute points and substantially shrinking the gap with strong proprietary models. All D3-Gym artifacts (environments, creation workflow, trajectories, and models) can be found at https://github.com/OSU-NLP-Group/D3-Gym.
翻译:尽管近期在面向科学数据驱动发现的语言模型和智能体方面取得了进展,但缺乏代表真实世界科学任务的可验证环境阻碍了其能力的进一步提升。为填补这一空白,我们提出了D3-Gym——首个为科学数据驱动发现而自动构建、包含可验证环境的数据集。D3-Gym包含:(1) 来自四个学科239个真实科学知识库的565个任务,其中 (2) 每个任务均配有自然语言指令、预装依赖项的可执行环境、输入数据集与工件预览、参考代码解决方案,以及自动合成的评估脚本。对D3-Gym中验证信号质量的严格评估表明,我们的评估脚本与人工标注黄金标准的一致性达到87.5%,且在领域特定评估逻辑上高度对齐,证明了其科学可靠性。此外,在D3-Gym抽取轨迹上训练后,不同规模的Qwen3模型在ScienceAgentBench上均获得了一致且显著的性能提升,其中Qwen3-32B提升了7.8个绝对百分点,大幅缩小了与强专有模型的差距。所有D3-Gym工件(环境、创建流程、轨迹及模型)均可访问 https://github.com/OSU-NLP-Group/D3-Gym。