Wind turbine power curve models translate ambient conditions into turbine power output. They are essential for energy yield prediction and turbine performance monitoring. In recent years, data-driven machine learning methods have outperformed parametric, physics-informed approaches. However, they are often criticised for being opaque "black boxes" which raises concerns regarding their robustness in non-stationary environments, such as faced by wind turbines. We, therefore, introduce an explainable artificial intelligence (XAI) framework to investigate and validate strategies learned by data-driven power curve models from operational SCADA data. It combines domain-specific considerations with Shapley Values and the latest findings from XAI for regression. Our results suggest, that learned strategies can be better indicators for model robustness than validation or test set errors. Moreover, we observe that highly complex, state-of-the-art ML models are prone to learn physically implausible strategies. Consequently, we compare several measures to ensure physically reasonable model behaviour. Lastly, we propose the utilization of XAI in the context of wind turbine performance monitoring, by disentangling environmental and technical effects that cause deviations from an expected turbine output. We hope, our work can guide domain experts towards training and selecting more transparent and robust data-driven wind turbine power curve models.
翻译:风力发电机功率曲线模型将环境条件转化为涡轮机功率输出,这对于能量产量预测和涡轮机性能监测至关重要。近年来,数据驱动的机器学习方法已超越参数化、基于物理信息的方法,但因其不透明的“黑箱”特性而常受批评,这引发了对其在非平稳环境(如风力发电机面临的环境)中稳健性的担忧。因此,我们引入一个可解释人工智能(XAI)框架,用于从运行SCADA数据中研究并验证数据驱动功率曲线模型学到的策略。该框架结合领域特定考量、沙普利值(Shapley Values)以及XAI在回归问题中的最新成果。结果表明,学到的策略比验证集或测试集误差更能指示模型稳健性。此外,我们观察到高度复杂的最先进机器学习模型倾向于学习物理上不合理的策略。为此,我们比较了多种确保模型行为物理合理性的方法。最后,我们建议在风力发电机性能监测中利用XAI,通过分离导致预期输出偏差的环境因素与技术因素。希望我们的工作能指导领域专家训练和选择更透明、稳健的数据驱动风力发电机功率曲线模型。