The central problem we address in this work is estimation of the parameter support set S, the set of indices corresponding to nonzero parameters, in the context of a sparse parametric likelihood model for count-valued multivariate time series. We develop a computationally-intensive algorithm that performs the estimation by aggregating support sets obtained by applying the LASSO to data subsamples. Our approach is to identify several well-fitting candidate models and estimate S by the most frequently-used parameters, thus \textit{aggregating} candidate models rather than selecting a single candidate deemed optimal in some sense. While our method is broadly applicable to any selection problem, we focus on the generalized vector autoregressive model class, and in particular the Poisson case, due to (i) the difficulty of the support estimation problem due to complex dependence in the data, (ii) recent work applying the LASSO in this context, and (iii) interesting applications in network recovery from discrete multivariate time series. We establish benchmark methods based on the LASSO and present empirical results demonstrating the superior performance of our method. Additionally, we present an application estimating ecological interaction networks from paleoclimatology data.
翻译:本研究核心问题是针对计数值多元时间序列的稀疏参数似然模型,估计参数支持集S(即对应非零参数的指标集)。我们开发了一种计算密集型算法,通过对数据子样本应用LASSO获得的支持集进行聚合来完成估计。我们的方法是通过识别多个拟合良好的候选模型,并利用最频繁使用的参数来估计S,即通过"聚合"候选模型而非选择某个意义上最优的单一模型。虽然该方法可广泛适用于各类选择问题,但我们聚焦于广义向量自回归模型类(特别是泊松情形),原因在于:(i) 数据复杂依赖性导致支持集估计的困难性;(ii) 近期在该场景下应用LASSO的研究进展;(iii) 从离散多元时间序列重建网络的有趣应用。我们建立了基于LASSO的基准方法,并通过实证结果展示了本方法的优越性能。此外,我们还展示了一个基于古气候数据估计生态互作网络的应用案例。