We introduce GRACE-DS, a Guarded Reward-guided Agent Correction Environment in Data Science for pre-deployment evaluation of LLM-powered AutoML agents. GRACE-DS is a set of evaluation metrics in an isolated environment that can be applied to tabular ML tasks specific to a particular organization. It exposes agents to realistic workflow stages, from planning and data inspection through feature engineering, model development, validation, and code repair to final submission, while hidden executable validators measure not only final predictive performance but also leakage avoidance, reproducibility, protocol validity, correction behavior, and reward alignment. The strongest structured regime, flexible iterative interaction (our approach), achieves higher end-to-end normalized hidden-test quality than single-shot generation, unstructured interaction, and restart-based baselines, while also improving protocol-valid completion. Validated across more than 7,000 episodes, these results establish GRACE-DS as a robust platform for assessing the capacity of LLM-based AutoML agents to execute machine learning workflows under production-like conditions and in accordance with organization-specific requirements.
翻译:我们提出GRACE-DS,一种用于数据科学领域的受保护奖励引导代理纠错环境,旨在对基于大语言模型的自动化机器学习代理进行部署前评估。GRACE-DS提供了一组在隔离环境中运行的评估指标,可应用于特定组织的表格型机器学习任务。该环境使代理能够经历从规划与数据检查、特征工程、模型开发、验证、代码修复到最终提交的完整工作流阶段,同时通过隐藏的可执行验证器不仅衡量最终预测性能,还评估泄露规避、可复现性、协议有效性、纠错行为及奖励对齐性。在最强结构化框架下,灵活迭代交互(我们的方法)在端到端归一化隐藏测试质量上优于单次生成、非结构化交互及基于重试的基线方法,同时提升了协议有效完成率。经超过7000次试验验证,这些结果表明GRACE-DS可作为评估基于大语言模型的自动化机器学习代理在生产级条件下执行机器学习工作流并满足特定组织需求的鲁棒平台。