Data-driven building energy prediction is an integral part of the process for measurement and verification, building benchmarking, and building-to-grid interaction. The ASHRAE Great Energy Predictor III (GEPIII) machine learning competition used an extensive meter data set to crowdsource the most accurate machine learning workflow for whole building energy prediction. A significant component of the winning solutions was the pre-processing phase to remove anomalous training data. Contemporary pre-processing methods focus on filtering statistical threshold values or deep learning methods requiring training data and multiple hyper-parameters. A recent method named ALDI (Automated Load profile Discord Identification) managed to identify these discords using matrix profile, but the technique still requires user-defined parameters. We develop ALDI++, a method based on the previous work that bypasses user-defined parameters and takes advantage of discord similarity. We evaluate ALDI++ against a statistical threshold, variational auto-encoder, and the original ALDI as baselines in classifying discords and energy forecasting scenarios. Our results demonstrate that while the classification performance improvement over the original method is marginal, ALDI++ helps achieve the best forecasting error improving 6% over the winning's team approach with six times less computation time.
翻译:数据驱动的建筑能耗预测是测量与验证、建筑标杆管理以及建筑与电网交互过程中的核心组成部分。ASHRAE第三届建筑能耗预测大赛(GEPIII)利用大规模计量数据集,通过众包方式筛选出针对整栋建筑能耗预测的最优机器学习工作流。获胜方案的一个关键环节是数据预处理阶段,该阶段旨在剔除异常训练数据。当前预处理方法主要聚焦于统计阈值过滤或依赖训练数据与多超参数的深度学习方法。最新提出的ALDI方法(自动化负荷曲线离群值识别)虽能利用矩阵轮廓识别这类离群值,但仍需用户定义参数。我们在前期工作基础上提出ALDI++方法,该方法无需用户定义参数,并利用离群值相似性机制。我们以统计阈值法、变分自编码器及原始ALDI方法为基线,从离群值分类与能耗预测场景两个维度评估ALDI++性能。结果表明:尽管相较于原始方法,ALDI++的分类性能提升幅度有限,但其能实现最优预测误差——较冠军团队方法误差降低6%,且计算耗时减少六倍。