Missing time-series data is a prevalent practical problem. Imputation methods in time-series data often are applied to the full panel data with the purpose of training a model for a downstream out-of-sample task. For example, in finance, imputation of missing returns may be applied prior to training a portfolio optimization model. Unfortunately, this practice may result in a look-ahead-bias in the future performance on the downstream task. There is an inherent trade-off between the look-ahead-bias of using the full data set for imputation and the larger variance in the imputation from using only the training data. By connecting layers of information revealed in time, we propose a Bayesian posterior consensus distribution which optimally controls the variance and look-ahead-bias trade-off in the imputation. We demonstrate the benefit of our methodology both in synthetic and real financial data.
翻译:缺失时间序列数据是一个普遍存在的实际问题。时间序列数据中的填补方法通常应用于全面板数据,其目的是为下游样本外任务训练模型。例如,在金融领域,缺失收益率的填补可能发生在训练投资组合优化模型之前。不幸的是,这种做法可能导致下游任务未来表现中的前瞻偏差。使用完整数据集进行填补所带来的前瞻偏差与仅使用训练数据进行填补所产生的较大方差之间存在固有权衡。通过连接时间上逐层揭示的信息,我们提出了一种贝叶斯后验共识分布,该分布能够最优地控制填补中的方差与前瞻偏差权衡。我们通过在合成数据和真实金融数据中展示了我们方法的优势。