The simulation of fluid flows is computationally expensive due to the complexity of its governing partial differential equations. Machine learning models offer a potential surrogate, enabling learning from simulations and significantly faster predictions of flow fields. However, these models require large training datasets, which introduces a trade-off between dataset generation cost and predictive accuracy. In this work, we investigate the relationship between the size of the training-set and accuracy of the prediction when learning steady flow fields in an industrial-scale stirred vessel. A data set of steady flows is generated using Reynolds Averaged Navier Stokes (RANS) simulations in a range of realistic operating conditions, including impeller speeds and liquid heights. We train implicit neural representations of flow fields and compare purely data-driven and constrained variants. Model performance is evaluated using global mean squared error (MSE), qualitative spatial comparisons of predicted and reference flow fields, and tracer transport simulations. We find that the prediction error decreases monotonically with increasing training data, but also that it exhibits clear diminishing returns beyond moderate dataset sizes. Physics-based constraints significantly improve accuracy and reduce variability across training runs in low-data regimes, and they lead to more stable tracer-transport behavior. Furthermore, reasonable interpolation can be achieved over different impeller speeds and liquid heights. However, these benefits come with an increase in the complexity of training, and their relative advantage diminishes as the training set grows.
翻译:流体模拟因其控制偏微分方程的复杂性而计算成本高昂。机器学习模型可作为潜在替代方法,通过从模拟中学习实现流场预测的大幅加速。然而,此类模型需要大量训练数据集,导致数据集生成成本与预测精度之间存在权衡。本研究探讨了在工业规模搅拌罐中学习稳态流场时,训练集规模与预测精度之间的关系。我们采用雷诺平均Navier-Stokes(RANS)模拟生成一系列实际工况下的稳态流场数据集,涵盖不同叶轮转速与液位高度。通过训练流场的隐式神经表征,对比了纯数据驱动与约束模型两种方案。基于全局均方误差(MSE)、预测与参考流场的定性空间比较以及示踪输运模拟对模型性能进行评估。研究发现:预测误差随训练数据增加单调递减,但超过中等规模数据集后呈现明显边际递减效应。在小数据场景下,物理约束显著提升精度并降低不同训练轮次间的变异性,同时带来更稳定的示踪输运行为。此外,模型可在不同叶轮转速与液位高度间实现合理插值。然而,这些优势伴随训练复杂度增加,且其相对优势随训练集扩大而减弱。