Despite their growing popularity, data-driven models of real-world dynamical systems require lots of data. However, due to sensing limitations as well as privacy concerns, this data is not always available, especially in domains such as energy. Pre-trained models using data gathered in similar contexts have shown enormous potential in addressing these concerns: they can improve predictive accuracy at a much lower observational data expense. Theoretically, due to the risk posed by negative transfer, this improvement is however neither uniform for all agents nor is it guaranteed. In this paper, using data from several distributed energy resources, we investigate and report preliminary findings on several key questions in this regard. First, we evaluate the improvement in predictive accuracy due to pre-trained models, both with and without fine-tuning. Subsequently, we consider the question of fairness: do pre-trained models create equal improvements for heterogeneous agents, and how does this translate to downstream utility? Answering these questions can help enable improvements in the creation, fine-tuning, and adoption of such pre-trained models.
翻译:尽管数据驱动的真实动态系统模型日益普及,但其构建仍需大量数据。然而受限于传感设备及隐私问题,尤其在能源领域,这类数据并非总能获取。利用相似场景采集数据预训练的模型在解决此类问题时展现出巨大潜力:它能以更少的观测数据代价提升预测准确性。理论上,由于负迁移风险的存在,这种提升对异构智能体而言既非均匀分布,也无法保证普适性。本文基于分布式能源系统数据,针对该领域的若干关键问题展开研究并报告初步发现:首先,我们评估了预训练模型在有/无微调情况下对预测精度的提升效果;其次探究公平性问题——预训练模型是否为异构智能体带来同等增益,以及这种增益如何传导至下游效用函数?回答这些问题将有助于推动此类预训练模型在创建、微调及部署等环节的优化改进。